Amazing how much the mainstream media are still obsessed with LLMs and factual accuracy. I prefer to see them as the worlds greatest innovation as lateral thinking machines ie pattern A applied over pattern B to give a credible pattern C.
LLMs are, and will continue to be, used for applications where factual accuracy is important. The decisions that LLM outputs prompt will soon be having big impacts in people's actual lives, if they aren't already.
The appearance of credibility of the output makes this much worse, since inevitably decisions will be deferred to LLMs that they are not sufficiently accurate enough for.
Real world industry wont stand for "hallucinated" outputs, unless there is more innovation around UI/UX on outputs. For example, no way lawyers/bankers/doctors are going to use LLM in their current forms and limitations if they can't trust the outputs.
Disagree. As a lawyer I use LLMs with RAG to help me surface information all the time. Often, this allows me to find niche case law that I just wouldn't have had the time to find on my own. However, I double-check everything, and read all the original sources.
LLMs are best treated as the AI equivalent of a human assistant who is knowledgeable and fast, but also inexperienced, and thus, prone to making mistakes. You won't throw out the work of such an assistant-it'll still save you hours of effort. However, you won't take the work at face value either.
I use LLMs for helping with coding it's proving invaluable. I see it very similarly, an inexperienced assistant with very broad knowledge. I also find that they don't make too many mistakes in tasks such as refactoring or finding bugs, it's when you ask it to just wholesale generate code for you that you hit problems. If I were to just take the code as is, not test it and not check I understand it then use it it's me that would be making the mistake not the LLM.
I have a friend in law school at the moment, and while he obviously can't use AI for school, multiple professors of his have recommended he get familiar with using them now so he'll be efficient at using them after passing the bar.
It's the trick where you take a question from a user, search for documents that match that question, stuff as many of the relevant chunks of content from those documents as you can into the prompt (usually 4,000 or 8,000 tokens, but Claude can go up to 100,000) and then say to the LLM "Based on this context, answer this question: QUESTION".
Real world industry already relies heavily on hallucinated outputs. They are called humans.
One of the main differences between people who see the astonishing value of LLMs right now and people who are some combo of skeptical, dismissive, and indignant, is the expectation that for something to be a valuable source of information, it has to be factually accurate every time.
The entirety of civilization has been built on the back of inaccurate sources of information.
That will never change and it absolutely can't, because (1) factual accuracy is not something that can be determined by consensus in a variety of significant cases, and (2) factual accuracy as a concept itself does not have a consensus definition, operationally or in the abstract.
Absolutely frustrating to see these topics addressed as if thousands of years of intense thinking around truth and factual accuracy has not taken place.
The results of those inquiries do not support the basic assumptions of these conversations (i.e. that factual accuracy is amenable to exhaustive algorithmic verification).
If you have employees, you also need to create transverse structures (make people work in teams, set check lists, QA, HR, accounting, create corporate charters on gender equality and many other topics, tell many "statements that ..." or "our mission is ...", create corporate culture, ) because you can't trust humans employees at 100%.
Actually some of banks biggest failures were when only a few and in some cases only one person was in charge.
> Real world industry wont stand for "hallucinated" outputs
Of course they will, if the other benefits are large enough. Checking factual accuracy of a large corpus can often be considerably simpler than generating the large corpus to begin with.
But how long will factual accuracy remain below human level? I expect the incorrect information issue to be a short term problem. That new GPT-4 based legal AI the BigLaw firms are signing up for already produces some documents with a lower error rate than a human lawyer.
Perfection isn't necessary, it only needs to be better than the average human. In a year or two I expect those kinks will be worked out for countless job tasks.
If I'm making a tool for doctors, I don't need to surface the exact scientific fact the LLM recalled from memory, I can design an interface that surfaces verbatim text from sources based on the LLM's understanding of the situation.
No serious product should be surfacing a ChatGPT style chat window to the user, it's a poor UX anyways with awful discoverability.
Remember the outcry against using Wikipedia as a source in school-work etc? Pretty convinced that the fear that "LLMs just make stuff up" will gradually go away as they get better.
Also, at the moment you should view e.g. ChatGPT as your autistic well-read friend who really do know quite a lot of things, but only approximately, and is prone to make things up instead of being found out for not knowing. It's OK to ask your autistic friend how electricity works, but don't expect him to correctly cite the titles and authors of relevant research papers. Just common sense stuff if you really think of the LLM as a brain of compressed information and not a search engine.
I think there is more innovation to happen with LLM output, for example leveraging the attention weights to give information on what the key events in history should be paid attention to given a sequence of current events, etc. More predictive information beyond just generated text responses ("completions").