For a long time after the internet arrived on the scene, a lot of online news stories would reference websites, papers, polls, etc. without linking to them. There are still news sources doing this today. Sometimes such articles interpret or place context around their hidden references, but a lot of the time they just summarize.
Giving someone the text output of a LLM is very similar to publishing a summary without links to the referenced material. When you were querying your LLM, you could have asked specific questions or asked for a custom focus or point of view. Your intended audience might have questions or different concerns, but they're unable to interact with your LLM. What you have delivered is static and unresponsive. It has all the disadvantages of being machine output without the advantage of being interactive, the way your LLM was for you.
It may have to wait until compute is cheap enough that tokens are essentially free, but we need a system to pass "hyperlinks" to LLM's primed with context, ready to be interactively queried on a chosen context. It's being overly generous to assume that people are putting even 300 bits into a LLM for every 1000 bits of regurgitated writing they try to pass off as their own. When people post LLM output as if it were their own, I have no choice but to assume they had zero knowledge of the subject, but this query taught them what they wanted to learn, and now they're sharing that. That's fine, but please pass an interactive LLM link rather than static text.
Once we have "hyperlinks" for LLM sessions, perhaps we can share LLM output a little more usefully and honestly.
I've seen professional journal pieces refer to science journal articles only to go and read the original article and find that it draws a different conclusion than what is implied by the journalist.
My father used to complain about that 40 years ago, though in his case he was reading newspapers rather than professional journals. But he'd point it out to me often enough that I started to see the pattern. Scientist publishes paper saying "We may have found evidence of X, which suggests the possibility that Y may also be occurring". Journalist: "Scientists find X which proves Y".
This has been happening for decades; I still see it happening today*. My cynical suspicion is that words like "maybe" and "suggests the possibility" don't sell enough papers.
* Worst offender I can remember was actually from the summary of a paper published on the research institution's own website, so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote". Summary said "Exposure to X can, on average, cause a 40% higher chance of Y" (where Y was a negative health outcome). I clicked through to the study and read it. Turned out the confidence interval on that chance of Y was so wide, all you could say with 95% confidence was that exposure to X could do anything from reduce your chance of Y by 5 percent, or increase it by 85 percent, or somewhere in between. They had averaged -5 and +85 to get the scarier-sounding 40% number that they published in the summary, but the truth would have been far closer to "this confidence interval is so wide that we really can't conclude anything from this data". But that wouldn't be nearly as likely to get them grants, so they tortured the data in their summary so that it would look better.
In most cases, they are only useful to other scientists in their field, which does mean that they are useless to the general person until they show up in the form of a new commodity.
There is also incentives to adapt the message to the outlet. If you send that data to peer review and say we found 40% increase, the reviewers would reject it, so they have to moderate themselves. But if you send a summary to the university’s outreach outlet saying that we found something or nothing we don’t know, then they would also reject putting it out. So even for the authors, the incentive is to send a careful conclusion to peer review and an overblown one to popsci.
Btw, I also think a 95% confidence interval is just the wrong statistic to look at given that data, and that they could probably have analyzed it better.
> Worst offender I can remember was actually from the summary of a paper published on the research institution's own website, so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote".
It's a similar thing. We live in the "attention economy", and research institutions - particularly after the US President openly went and had his minions cut funding to research purely on ideological reasons, but it's been a problem for decades - are just as susceptible to blow stuff out of proportion to make headlines and thus increase the chance someone might throw some money over the fence.
And media does the same, just to manufacture artificial debate. And so do politicians.
And frankly, I'm fed up with that, we will drive ourselves into a wall.
> so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote"
It's never the case that someone misunderstood what scientist wrote. Much like the scientific papers, news articles, including those reporting specifically on the discovery, have their own goals, and the paper being cited is used as evidence or argument for article's own "study". Except for press, the standard is rhetorical, not scientific, it's the conclusions and not the methods that are "pre-registered" at the start, and claims are defended by "hey it's just a point of view", not by statistical significance.
In your own example of worst offender: the scientific study was trying to establish and quantify the connection between X and Y. The summary article was trying to push the angle that "this institution is doing important work". It started with that conclusion, and the paper cited was just the first thing the author found that could be easily massaged into supporting that conclusions by rhetorical standards.
Same paper might get cited by journalist trying to push for "X is bad for you", and they'll do roughly the same as the summary article. And, same paper may be cited by someone claiming they have a miracle cure for Y, and they'll make a honest observation that "absence of X reducing Y is a common bullshit claim based on misunderstanding the paper [citation], that actually shows there's no correlation there, I mean look at the confidence intervals, even the author says that in text nobody bothers to read"... - citation may be honest, but the article itself is still using it to prop up a different flavor of bullshit.
TL;DR: don't believe news. It's bad for your mental and physical health (p<00.05).
> It's never the case that someone misunderstood what scientist wrote.
This is wrong. Reporters frequently don’t understand the science or the nuance in the science.
Reporting and science are two very different disciplines. Reporters rarely have a deep background in science and almost never have a background in the specific area that they’re reporting on.
Hell, even scientists have trouble accurately describing the work of a different scientific discipline.
Don’t invent bad faith motivations; they exist but most of the time it’s just two people slightly talking past each other.
> Don’t invent bad faith motivations; they exist but most of the time it’s just two people slightly talking past each other.
I'm not inventing them, but maybe conflating two sources:
1. Malice directly intending to hurt or defraud people. Probably not as common as how I make it seem.
2. Not caring. Well, I subscribe to the view that not expending effort to be accurate when talking to other person is as bad as slashing their tires (paraphrasing an old quip), so I very much consider bullshitting and picking a conclusion and then massaging facts to fit it, to be acting in bad faith too.
Nah, even in middle and high school it was plain that people sometimes misunderstood what the teachers meant. It's not like that goes away in adulthood; even when people "care", they still often misunderstand nuance or details, and sometimes even the bigger picture.
Honestly, it's plain weird to say that people never just make mistakes.
PS - worth adding that "I misunderstood" and "I didn't care enough" are not mutually exclusive. You can do both, so saying "they didn't misunderstand, they just didn't care" isn't a reasonable rebuttal. But even setting that aside, there'll be plenty of folks who care but still don't understand.
I'm gonna invoke Hanlon's handgun here: not attributing stupidity to what's adequately explained by systemic incentives promoting malice.
I normally assume people make mistakes. I don't believe this is a good explanation for news publications, university press releases, politicians, etc. because those are organizations with agenda, and commit "misunderstanding" of this type pretty much in every thing they publish. The pattern here is pretty conclusive, IMO.
> I normally assume people make mistakes. I don't believe this is a good explanation for news publications, university press releases, politicians, etc. because those are organizations with agenda, and commit "misunderstanding" of this type pretty much in every thing they publish. The pattern here is pretty conclusive, IMO.
The pattern of personal motives fits misunderstanding. You need to show that there is an organizational pattern of "malice" (your word, not mine), rather than an organizational pattern of "we are trying to publish quickly, and quality accidentally falls to the wayside". I.e., negligence, not malice.
You haven't provided even a shred of evidence suggesting there's malice at the journalist level. Every science journalist I have met genuinely cared about the science (which is why they were writing on it), but they didn't have time to learn enough about the subjects to understand they were oversimplifying things.
Not in science journalism, but I've personally encountered a case in normal journalism that I can only attribute to malice. It was many years ago, but it was so blatant I still remember it.
The 911 call went like this, according to its transcript. Caller: "This guy looks suspicious, like he's on drugs or something. It's raining and he's walking around looking into windows." 911 operator: "Can you describe him? What race is he?" Caller: "He looks black."
How the TV news reported it on the air was: Caller: "This guy looks suspicious ... He looks black."
Omitting excess verbiage is one thing. Omitting words that entirely change the context of the statement, making it look like the caller was racially prejudiced rather than responding to a specific question, is something else entirely. That was the last time I trusted reporting from that particular source (it was NBC, by the way).
My principle is that when someone lies to me, I stop trusting them. By lying I mean not just omitting details, or having an obvious bias, but deliberately telling me A when they clearly know that the truth is not-A. I could not see that report any other way but a deliberate lie, knowing the truth and attempting to make people believe the opposite.
I have recently reviewed a paper that referenced my own article… but the conclusion was so off, I actually went back and re-read that entire article just to be sure there was no hint to the conclusion that the author derived. There was none. Uncanny experience.
Link rot has accelerated in the past decade. Links are great, when they work. If they are a few years old, they often don't. You certainly can't count on it. If you're referencing a paper or an article, it's probably better to cite the title of the article, the name of the journal or magazine or newspaper it was published in, the author, and the date. Then someone might have a chance of finding it again.
Not just news stories. This is a huge problem with social media and forums in general too. Lots of people making bold claims about random things, lots of stories they say are based on a third party source, but very few actual links to said sources in question.
I still remember a recent example where one of those trivia accounts on Twitter posted an interesting story about some guy whose life completely changed after an accident, but neither linked to a source or named the person in question.
The only way I was able to verify it was true was through someone in the comments asking the platform's AI chatbot, and the chatbot providing context that I could research and verify...
> For a long time after the internet arrived on the scene, a lot of online news stories would reference websites, papers, polls, etc. without linking to them. There are still news sources doing this today.
I agree, this drives me crazy. Ironically, one of my favorite uses for Claude is to ask, "What study is this news article talking about?"
It's pretty good at digging up the source and related sources. And most of the time, if you read the source, the article is nonsense and gets everything wrong.
> For a long time...a lot of online news stories would reference websites, papers, polls, etc. without linking to them.
The "$CITY_NAME Business Journal" websites are the absolute worst with this. They'll refer to something specific, for example "$BIGCO's 2025 10-K filing" and it will be a link. That link will go to the 10-K, right? Nope! It goes to another page at the same business journal. Maybe that page is a summary of the 10-K, but probably not. Maybe it's just the general index page for all the articles about $BIGCO at that journal. What it links to, it definitely won't be the specific thing described by the text of that link.
Giving someone the text output of a LLM is very similar to publishing a summary without links to the referenced material. When you were querying your LLM, you could have asked specific questions or asked for a custom focus or point of view. Your intended audience might have questions or different concerns, but they're unable to interact with your LLM. What you have delivered is static and unresponsive. It has all the disadvantages of being machine output without the advantage of being interactive, the way your LLM was for you.
It may have to wait until compute is cheap enough that tokens are essentially free, but we need a system to pass "hyperlinks" to LLM's primed with context, ready to be interactively queried on a chosen context. It's being overly generous to assume that people are putting even 300 bits into a LLM for every 1000 bits of regurgitated writing they try to pass off as their own. When people post LLM output as if it were their own, I have no choice but to assume they had zero knowledge of the subject, but this query taught them what they wanted to learn, and now they're sharing that. That's fine, but please pass an interactive LLM link rather than static text.
Once we have "hyperlinks" for LLM sessions, perhaps we can share LLM output a little more usefully and honestly.