Typically issues arise because they are novel and are unforeseen. If we did see these issues beforehand they'd be fixed! LLMs by definition are trained by example, so I fail to see how finetuning LLMs on things that have already happened to be helpful for determining the root cause of a novel issue.
LLMs seem to lack systemic modelling that humans do. I can see LLMs being practical for this if it is shown that LLMs are capable of modelling scenarios outside of their dataset, but thus far none such examples exist.
During my time taking on-call pager rotations at fairly large engineering organizations, I'd say that a distressingly large number of incidents happened for fairly standard reasons ("didn't write tests for this case" + "deployed" + "insufficient monitoring," with things like incorrect concurrency assumptions around DB access, or memory leaks in application code, or insufficient rate limiting / circuitbreaking causing cascading failures being fairly common passengers). In fact, it's actually pretty hard for me to come up with an issue that wasn't basically a fairly common programming error combined with some relatively common infrastructure, build, test, and/or observability problem. Sometimes with some common bureaucratic human problems thrown into the mix too, i.e. "no one owns this service's uptime."
I think we're at an epistemic impasse here. At what point would/could you be convinced that LLMs are incapable or unsuited here? If LLMs were successfully deployed in a production environment is the day I bite my tongue. What about you?
I'm not even sure that LLMs are even capable of solving standard bugs see: [1]. Hallucination seems to be a significant hurdle and any time spent validating the fixes of an LLM is wasted when it could be spent tackling the bug head on. The amount of energy spent espousing garbage requires an order of magnitude more effort to invalidate.
Alright, fine. Maybe you don't have faith in management, but perhaps you do have faith in the open market and capitalism.
Feel free to point out any error in my logic:
There are huge financial incentives--tens if not hundreds of billions of dollars--for developing an LLM which can solve novel bugs. So surely there exists AI companies developing an LLM capable of doing so. If an LLM capable of solving novel bugs exists, AI companies would rush to showing it off to capture tonnes of VC money. AI companies could show off their fancy bug-fixing LLM by closing issues on public Github repos using said LLMs.
No such mythical LLM exists. We are thus left with two choices:
1. My logic is flawed or there is an alternative possibility I haven't considered.
2. The LLM capable of doing what OP asserts doesn't exist and can't be made, despite their assertion that it is trivial to fine tune and put into application.
The base technology capable of this has only been broadly available for about a month — prior to Llama-3.1-70b being released on July 23rd, you couldn't finetune any GPT-4 class models that had long context support (OpenAI only allowed fine-tuning their GPT-3.5-Turbo model until last week), and you'd need long context for the incident data — so I think "proof by inexistence" is pretty weak here. The first personal computer shipped in 1974, but it took five years until the development of Visicalc for spreadsheets to appear, despite the huge business value. I wouldn't expect most use cases for LLMs to appear within a month of them being made available.
To answer your question more directly, I would go with option 1.
It's not proof by inexistence, it's simply application of the scientific method--only the most successful method to date.
It seems like your assumptions are unfalsifiable. As computer scientists, I believe it's important that our hypotheses are testable. If a hypothesis is unfalsifiable, then the hypothesis is no better than theology and should be discarded.
What's stopping you and other VCs just pouring endless money into an idea that won't work?
The scientific method generally involves experiments, as opposed to claiming something won't work because of the "logic" that if it worked, someone would have done it already. This particular hypothesis is obviously a testable one: someone could simply follow the proposed steps from the hypothesis (e.g. finetune a model on their incident response data), and see if it works. This is in fact how essentially all machine learning research is done: coming up with a proposal and trying it out. If it works, great! You've probably contributed something new to machine learning research. If not, oh well, try and figure out why your experiment failed, and if you have a good alternative approach try that instead.
Your variant of the "scientific method" would've meant we never discovered electricity, or invented airplanes, or really anything else, because why bother trying? If it worked someone else would've done it.
> This particular hypothesis is obviously a testable one: someone could simply follow the proposed steps from the hypothesis (e.g. finetune a model on their incident response data), and see if it works.
Are you saying that if someone finetunes a current SOTA LLM with incident response data and demonstrates that it doesn't work that you'll say that LLMs are infeasible for this application? That would invalidate the hypothesis: "X application can be done on current LLMs."
Such a test could never invalidate the the hypothesis: "X application can (eventually) be done on LLMs."
If it's the former hypothesis you were asserting, then yes I agree that it is testable, but I'm fairly confident you were asserting the latter.
Earlier I had asked you: "I think we're at an epistemic impasse here. At what point would/could you be convinced that LLMs are incapable or unsuited here?"
It would invalidate that particular approach, much as a failed attempt at creating a lightbulb would invalidate that particular approach, but would not disprove the lightbulb entirely.
Proving that LLMs can never do this would require extremely rigorous theoretical evaluation that even top ML labs are currently unable to do, given the problem of interpretability. In general proving a negative is typically harder than a positive, since a single experiment succeeding proves a positive, but a single experiment failing does not prove a negative; generally science does not demand that scientists attempt to prove a negative when running experiments, or else nearly every drug trial, for example, would be impossible to perform. Complaining that you have staked out a very difficult to defend position — that it's impossible for LLMs to generate good incident reports — does not mean your ideological opponents, who have simpler positions, must do your proof work for you.
Just to make sure we both understand what burden of proof is:
Suppose two people are having a debate over whether or not a teapot exists in the orbit of Jupiter which is impossible to observe via telescope. Where does the burden of proof lie?
Just to reiterate plainly:
Does the burden of proof lie on the person making empirically impossible to falsify claim or the person making the empirically possible to falsify claim?
Which of the following two claims is impossible to empirically falsify?
1. "LLMs can eventually be used to produce good incident reports."
2. "LLMs can never eventually be used to produce good incident reports."
Sorry, but saying "I bet this would work" does not mean I have to come up with a theoretical foundation for disproving the existence of machine learning models ever being capable of doing things, theoretical models which even the top labs in the world are incapable of producing. This is the hypothesis stage; there is no burden of proof. If I said "I proved this would work," naturally there would be a burden of evidence. That is not where we are. You are arguing with a hypothesis; and your argument does not hold water ("If this was possible someone would have already done it"). That does not mean the hypothesis is true, it only means you haven't falsified it.
So at this point are you admitting fully that you have no evidence for your claim? Your entire bet is conjecture?
> That does not mean the hypothesis is true, it only means you haven't falsified it.
Correct. You can't prove a negative, you've only just figured this out? After I listed TWO simple examples that could be found in an introductory philosophy class?! In the Wikipedia article consisting of a couple of paragraphs I linked you?
PLEASE JUST READ. PLEASE JUST READ. PLEASE JUST READ. PLEASE JUST READ.
You can't prove negative statements. I want you to admit this so I know you understand, now repeat after me: "You can't prove negative statements."
I know this is likely wasted on you, but here goes:
You can't prove the hypothesis: "There does not exist an unobservable teapot in the orbit of Jupiter."
For the SAME reason I can't prove the hypothesis: "LLMs can never eventually be used for generating good incident reports."
For the SAME reason I can't prove the hypothesis: "God doesn't exist."
For the SAME reason I can't prove the hypothesis: "A unicorn does not exist at the center of the Earth."
I fully admit this in the comment you've supposedly read.
This is why the burden of proof lies on the person making the positive claim (this is you).
We're literally discussing a paragraph in which I said "I'd bet [this would work]." No amount of repeatedly claiming there's a "burden of proof" to a bet and all-caps shouting and demanding that I come up with a proof of impossibility that would show the hypothesis is wrong (you may also note that I mentioned that proving a negative is generally more difficult than proving a positive several posts ago — although, in fact, it is technically possible [1]) will make your position a reasonable one. Politely, I will not be continuing to engage with you.
My point was that management has a history of rolling out shiny things to production and then having egg on their face. See Microsoft's racist bot, Google's AI making up stuff in their adverts, etc.
Your original wager was that it would be in production, not that it would work.
It's a funny thing, because the inverse is falsifiable (testable) whereas the positive version is not. The inverse proof (I would say a hypothesis) is simply application of the scientific method.
There is a way to disprove that the statement: "There is no god" by simply showing a counterfactual god.
There is however no way to disprove the statement: "There is a god."
Likewise, there is a way to disprove the statement: "LLMs cannot be successfully used for X application." By showing that LLMs have been used in X application.
Again, there is no way to disprove the statement: "LLMs can (eventually) be used in X application."
The meat of my question was meant to demonstrate a failure to apply the scientific method.
>There is a way to disprove that the statement: "There is no god" by simply showing a counterfactual god.
that both parties to the argument agree is a god.
>Likewise, there is a way to disprove the statement: "LLMs cannot be successfully used for X application." By showing that LLMs have been used in X application.
Again there the point of argumentation will be the word "successfully", the LLM would have to be such an overwhelming success at what it is trying to do that one cannot weasel out of it with "successfully".
Maybe. But in the list of lucrative applications I think bug-fixing is near the top. I think it's lucrative enough to attract at least a decent chunk of engineering talent.
Agreed. There have been many assessments of what bugs cost, and the assessments are often very high, and that's the reason the industry has, for decades, been working towards having _fewer_ bugs.
Typically issues arise because they are novel and are unforeseen. If we did see these issues beforehand they'd be fixed! LLMs by definition are trained by example, so I fail to see how finetuning LLMs on things that have already happened to be helpful for determining the root cause of a novel issue.
LLMs seem to lack systemic modelling that humans do. I can see LLMs being practical for this if it is shown that LLMs are capable of modelling scenarios outside of their dataset, but thus far none such examples exist.