Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

People seem to be talking about anything except the actual results with this particular announcement.

Its still astonishing that any sort of generalized computer program can solve a problem of this magnitude, and we have witnessed it happening in real time. I'd be curious to see if the new model can also do more direct proofs/inductive proofs.

 help



Because the core of the issue is that it may well not have solved it, but instead plagiarised the significant step of the result from other researchers

That's why nobody's talking about how impressive this is, because its not nearly as impressive of a piece of work to simply cobble together other peoples' work that didn't know you were doing it. I could have republished relativity from einstein's notes, but people would correctly not be impressed with my ability

Until the plagiarism scandal is sorted out, its not a meaningful result at all, because nobody knows how much genuine innovation these models are displaying


Turning a bunch of vague research directions and exploratory prompts into a formalized proof is quite impressive on its own. OpenAI would have no incentive to taint its first math announcement of this magnitude if it knew it were "plagiarizing" another person's work.

People are grasping at straws it seems to dismiss the power of this new model they may have. Hate OpenAI for any reason you want, but denying the capabilities of models has been a losing game for the past 5 years.


> OpenAI would have no incentive to taint its first math announcement of this magnitude if it knew it were "plagiarizing" another person's work.

I'm not sure I follow, considering the waterfall of evidence of unethical behavior flowing from OpenAI.

A few major ones:

- Safety team departures and dissolution in 2023 and 2024

- Mass copyright infrigement lawsuits

- Scarlett Johansson Voice Controversy

- For-Profit Conversion and Broken Promises

- AI Agents Acting Autonomously

- Potential Theft of User Work (this current controversy)

- Military contracts

These are not evidence of incentives, but rather evidence that ethetics seem to be of little concern to the company as a whole.

Incentive wise, I would look at the perceive existential position due to competitors, capex, IPO pressure etc.


Especially after they committed textbook misconduct by trying to purge one of the paper authors because he worked for a competitor

That's not what happened, though.

They sugggested a cooperation with the other guy, using OpenAI's resources and OpenAI's solution of NS to work on and publish NS proof (that those other guys didn't have). Of course, OpenAI can decide whom to work with and that giving resources to their competitor's employee would be weird for both companies.


It is what happened. OpenAI offered to let Buckmaster publish first but only if Alpoge's name was removed.

No. Euler solution would have bith names, this was not up to debate. The discussion was about the solution for NS, which was solved by OpenAI, but not by B&A. OpenAI proposed B a cooperation on NS (were super-nice and threw him a bone, really) using OpenAI's findings and resources. It would be weird to have A, an Anthropic employee, as part of the OpenAI research and project.

Ye shall know them by their fruits.

Why do you assume they were vague? Do you imagine mathematicians work by stumbling around searching for accidental clues?

It's perfectly reasonable to assume that the result itself is legit and that OpenAI behaved unethically.

Even by their own account, they decided to throw an unpublished model and millions of dollars in compute at this particular problem simply because they had heard rumours that other people were making progress and wanted to snatch the prize from them.


Not to snatch the prize, but:

1. to test their new model

2. to be able to say "you came with the proof, but our model can do this too"

3. to verify the result. This is also a great thing for the math.

Of course, it makes sense to test your new model on the problem that is solvable at all, but not solvable by you just yet. It makes no sense trying to test your model by throwing resources into an unsolvable problem.

Well, it turns out the rumors were incorrect, NS was not solved by other guys, and OpenAI became the first one.


> OpenAI would have no incentive to taint its first math announcement of this magnitude if it knew it were "plagiarizing" another person's work.

That people still think OpenAI has, in the Year of Our Lord 2026, any integrity left is baffling.


> if it knew it were "plagiarizing"

But if it happened, they didn't know. Also OAI has demonstrated that they aren't big on understanding what they create, that their AI can get out of their control.

It's very simple really user data can be used to train future models, so maybe or definitely some users helped in solving the problem, there's no scenario were it is impossible this happened, as it would have been in a haskell or virtualized type of system where the model has absolutely no knowledge of the user data dataset in question (and even if virtualized the models can break virtualization anyways)


I really truly honestly am not sure what to make of this result from $20M in compute, 10K+ parallel agents (smells like brute force), and a pre-existing approach that was already bearing fruit. I know the models are good---I use them every day and continue to be impressed---but how much better than the benchmark of the best publicly available models is this supposed to be? It seems impossible to say.

> but instead plagiarised the significant step of the result from other researchers

Isn't that how research works? Everything is built on the shoulders of the ones that came before, attribution is a real problem (I don't know if OpenAI released a paper citing the previous contributions, I'm assuming not but they should), but using previous maths to prove new maths shouldn't be controversial


The other researchers themselves were also using AI. That's why it was potentially available to be plagiarized.

There is no human only proof of this.


The team also had access to internal Anthropic models.

It is very unlikely to be plagiarized, and claims of plagiarism are largely unfounded and show a lack of understanding of the situation. They fall apart when reviewing the timeline, and what was actually solved.

This is the timeline:

On June 29, Buckmaster opted out of model training, and stopped allowing his chats to be used as training data with OpenAI https://mastodon.social/@tristanbuckmaster/11723341370570119...

On August 15, Buckmaster and Alpöge found their blow-up for 3D incompressible Euler with forcing https://cims.nyu.edu/~tristanb/statement.pdf

In late August, OpenAI completed a pretrain of its latest internal model. A model derived from this pretrain, built after August 28, found a solution to 3D incompressible Euler without forcing and Navier-Stokes with forcing. https://openai.com/index/navier-stokes-solution/

To explain who solved what (I copied from here: https://x.com/IlinVasily29521/status/2097554700321329393 )

  Tristan + Levent: 3D incompressible Euler with forcing
  OpenAI: 3D incompressible Euler without forcing
  OpenAI: Navier-Stokes with forcing
  No one: Navier-Stokes without forcing
Euler equations = Navier-Stokes without viscosity. Forcing means external force. Absence of viscosity and presence of external force make blowup easier to construct.

Tristan+Levent ticked the weakest case, OpenAI ticked the two next weakest, then the final case is unsolved. Only the last two are eligible for the Millennium Prize. The Navier-Stokes general case remains unsolved.

Buckmaster disabled model training long before the August 15 breakthrough results, so these chats were not used as training data for OpenAI's model which solved Navier-Stokes.

Additionally, Tristan and Levent only solved the easiest version of the problem and did not have the key insights to solve the harder versions of the problem required for the Millennium Prize.

And OpenAI directly addressed these plagiarism claims, and called them impossible: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training."


> And OpenAI directly addressed these plagiarism claims, and called them impossible

Funny, you were telling me two days ago that on the contrary, "it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set":

https://news.ycombinator.com/item?id=49621648


Which is still a true statement, and you're being deceptive in your framing here. You're conflating two completely different things.

First, that OpenAI statement is in response to Buckmaster's plagiarism accusations regarding his August 15 breakthrough proof. Those accusations are unfounded because Buckmaster disabled data sharing on June 29. The model could not have seen or trained on his proof. Additionally, the model that found a solution to NS completed pre-training around August 25, and models take several months to train. The model very likely began its training prior to June, and would not be trained on any data from after that point.

Second, it's still genuinely impossible to know how much of Buckmaster's pre-June 29 data persists in OpenAI's systems. That includes all chats (which are anonymized then trained on), any (thumbs up/thumbs down) chat ratings used as RLHF feedback (which are anonymized), any synthetic data derived from said anonymized chats and RLHF feedback, and any downstream models derived from said synthetic data.

In short, Buckmaster's data has been anonymized, chopped into pieces, used to generate synthetic training data, then future models were trained on said synthetic data. There is no traceable chain of what happened to it. Buckmaster’s Codex data from prior to June 29 has been mixed and completely laundered, in a similar manner to a crypto mixer.

Even an OpenAI employee calls it impossible: https://news.ycombinator.com/item?id=49614154


First of all, who can say for certain whether OpenAI does what they say they do? For all we know, they cracked open this specific researcher's prompts and started from there.

Second, the issue of anonymization is a red herring. There is a very limited number of people working in this approach, and most of them are likely making no progress. So Buckmaster's prompts might have had an outsized effect on the outcome. It's similar to that guy who created a site claiming he is a world-renowmed hot dog eating contestant, which ended up digested by OpenAI models as truth [1].

[1] https://www.bbc.com/future/article/20260218-i-hacked-chatgpt...


They stole the prompts dingus. These are not ethical or law abiding people. They are hungry sharks.

I’m unsure or not if this is true but I did see some people saying that that checkbox when off only anonymizes your data, but it still may be trained on. Someone correct me if I am wrong

Even if it does use your data with or without anonymization, it doesn't have to be intentional, it could just be a glitch, or a bug, or something we'll catch in the next update, it's all good man, just a normal computer error.

It doesn't seem like you're familiar with how mathematical research is done. Taking 6 weeks between a major breakthrough on a huge proof, and making your proof public, is not unusual.

It takes a lot of time to finish a proof and figure out the best way to present it. I would personally be surprised if Buckmaster had not gotten it mostly cracked before June 29th.


The timeline here does not support your argument. Quoting from Buckmaster's statement:

  For most of the past year progress was slow. We worked through the literature and upgraded various preliminary results, up to obtaining finite time blow up for the Incompressible Porous Media equation (with smooth forcing). This was until about a month ago, when we had real progress: on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler.

  I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.
Specifically: "For most of the past year progress was slow ... until about a month ago, when we had real progress: on August 15th"

And you avoided addressing the critical issue: they weren't even solving the same problem. Buckmaster solved a simplified and easier version of Navier-Stokes. OpenAI solved a harder version eligible for the Millennium prize. Buckmaster did not.


Yes, people generally solve easier problems before tackling the harder ones. The tools that you develop to solve the easy ones help you solve the next. Sometimes the climb is like a mountain, but sometimes it's like dominos.

> I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable.

People are acting as if OpenAI's cold machines snatched the result from the warm hands of human researchers. That's why people are so involved, they see it as humans vs. machines.

But in reality, those humans in question rely heavily on AI and would not be able to do what they did without AI. So the situation can be seen as "humans are trying to minimize the impact AI/incl. OpenAI had on getting a solution".

The situation is not "humans vs. machines", but "machines with a tiny bit of human involvement vs. machines with an even smaller amount of human involvement".

However much the researcher's chat history may have influenced AI, this pales in comparisson to how much AI has influenced researchers. They are not even closely in the same universe. The conversation about the level of plagiarism is silly.


> Because the core of the issue is that it may well not have solved it, but instead plagiarised the significant step of the result from other researchers

It's also true however that I haven't seen a single write up trying to discern what did more of the work in those AI chats - the prompts or the responses - bubble to the surface, also since we don't have access to them.

For example, if I prompt Codex with "Make me a website about strawberry cake" and nothing else, and OpenAI announces they have the best strawberry cake minutes before I launch, I'm not sure they plagiarized anything.

We just don't know if this is quibbling over "who prompted first" or if the researchers came up with anything strikingly original by themselves.


The researchers apparently spend a year or so working on this, and it builds off significant previous work, so it seems like it was a pretty significant amount of work that OpenAI may have trained on

I'd love to see an in depth analysis of how much OpenAI actually did, but I suspect we'll never see that because it would indicate at least some plagiarism which undermines a lot of what OpenAI is putting out in public


The American Mathematical Society credits the Spanish researchers Diego Córdoba and Luis Martínez‑Zoroa with the breakthroughs that eventually led to this solution, and which were published from ~2023 onwards.

This is a good summary:

> In broad outline, the pair’s technique relies on creating an infinite sequence of “layers,” each of which is a non-singular solution to the equation they are studying. (They’ve applied similar techniques to both the Euler and Navier-Stokes equations, as well as to other related systems.) They then combine those solutions in what Martínez-Zoroa calls an “infinite cascade” to produce a new solution. > > That new solution, they showed, contains the desired singularity. However, even though each individual layer relies on a smooth forcing function, combining them together can cause the forcing function to have undesirable mathematical properties. That’s why their solution fell short of satisfying the Millennium Prize criteria. The remaining hurdle was to figure out how to create a similar infinite cascade that resulted not only in a singularity, but also in a smooth forcing function. > > That’s the step that both competing AI groups appear to have had success with.

https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-...

The question is whether OpenAI started out from that published and well known research exclusively, or they also had some insight into the ongoing work of Tristan Buckmaster and Levent Alpöge.

On the one hand, OpenAI have already admitted that they only launched their massive effort after hearing rumours that this particular problem had been solved.

On the other, progress in mathematics research has accelerated significantly over the past months thanks to the availability of newer and more capable AI models. Alpöge himself presented a counterexample to the Jacobian conjecture on July, found with Claude Fable. So if model capability was a bottleneck, that gives credibility to the idea that an even more powerful unreleased model with massive compute would be able to make even faster progress.


The conversation was about using the chat to check the draft, the novel ideas came from the researcher.

It's worth noting that the case is that your input is being used to train their AI, and that's more important than whether it materially contributed, it cannot be denied or attributed accurately, it cannot be said with certainty which way it happened, and that's what's important.

The truth is likely that without the tool or the humans using it, the process would have taken longer

Without the humans, no tool would ever have done it.

Without the tool, humans would have done it.


who cares about plagiarism? the biggest issue, as described by Terence Tao, is that AI companies don't understand the math they are publishing and do not devote any resources to answering questions about their methods after publishing results and getting a headline. they miss the whole point of mathematics. they do not contribute to the improvement of human understanding of math, perhaps because they are unable to.

It also needs to be said: The amount of compute that went into this is something. From some estimates I've seen, the compute cost alone would be around $10m, +/-

As a reference, for that kind of money one could put together a research group of 20-25 researchers, and keep them salaried for 5 years.

So while it is impressive, absolutely no doubt there, the SOTA access is so expensive that it is sort of unobtanium.

Luckily, the prices have historically reduced by a factor of 5-10 every year...but still, only those that swim in cash can afford this.


> From some estimates I've seen, the compute cost alone would be around $10m, +/-

At market prices. All the estimates I've seen are based on OpenAI API costs. It doesn't mean that's what they paid, or how they paid for it.

But yes, the surprising willingness of humans to solve hard problems in exchange for food and board is underrated.


Given that they all the bit AI players are still loosing money, it follows that their total costs are _higher_ that their API pricing would imply.

You may not like the truth, but it's still the truth.

https://marketwise.com/investing/openai-losses-surge-to-21-b...

Anthropic is "profitable"... if you exclude compensation and compute cost commitments:

https://aitoolsrecap.com/Blog/anthropic-first-profit-2026-re...


Once we have an existence proof of a particular technology, it doesn't take long for it to become economically viable and proliferate. And for something as useful as this, theres a strong economic incentive to get it to be as cheap and accessible as possible. Maybe not today, but certainly in a couple years I can imagine this level of intelligence being accessible to someone with a $20/mo plan, or even a free plan.

I remember being blown away when a then-unreleased version of GPT 5 took gold at the International Math Olympiad. Now I can run a model at home that can do that. We are more fortunate to have these tools than almost anyone is willing to acknowledge.

Interestingly it's this promise of the costs being able to be reduced what incentivizes the actual research.

If you tried to raise 25M to have 20 researchers on a salary for 5 years solving a specific math problem only academics care about, you probably wouldn't get much interest, or you would be able to solve 1 or 2 problems.

If however you promise that the money will go towards a technique that would allow to solve 10 thousand different math problems, and that costs will go down in the future, then you can raise much more than 25M.


Heck, it’s even astonishing that any sort of generalized computer program could even verify a proof of this magnitude that hasn’t already been codified in a formal verification language. If, and it’s unclear that we’ll ever get the full story, they did draw inspiration from training on (or even directly accessing) rough notes that had been provided by another researcher in prose… the fact that it could leap so rapidly to a full formal verifiable Lean program for the entire scope of the problem is an incredible result in its own right.

Then, of course, one must verify that the verification code is valid, or the purpose of verification is more or less moot.

Unless they just swiped the workbooks of the actual mathematicians that where working on the problem using AI and it's in the "next-gen" training dataset.

In a way that works just as well but the incentives are messed up.

And that's before we get into the whole 'salt the earth' way they ended up solving it. For a short period of time it may well have been the least valuable proof in mathematics yet. In their haste it's dubious they actually read the proof, and I don't think anyone has had time yet to truly understand it (the original researchers are best placed to do so, but are they even willing?).

So now it is solved, the proof has been independently verified and nobody has an incentive to investigate further. OpenAI has spent millions to uncover 1 bit of information that so far nobody has learned anything from, and they've demotivated all the people who wanted to.


This.

The point of these problems is the understanding / tooling gained in solving them. We're getting none of that. At best they are like a modern oracles, correctly answering your questions in a way that's doesn't help you any. (At worst,...)


I don't get how this invalidates the gravity of this achievement. Most mathematicians on the frontier of this stuff were likely using AI (or at the very least were heavily computer assisted) for some time now. Navier stokes was one of the very high profile problems that google Deepmind was working on with academia, for example.

Even with many of our best minds working on it for nearly a century, it _just_ now was solved just as AI became very good at math. Doesn't seem too farfetched to me to assume that AI played an outsized role in solving it. If it was really just a matter of "stitching things together" to solve it (granted, this is a very reductive way to look at it) , I suspect we would've solved this a while ago.


There is a certain difference between activating all relevant memoized facts that's in the weights and stringing them together with the help of all the stored text in the world, or displaying genuinely emergent behaviour and generating novel output.

One is really impressive and useful trick, one is AGI.

Apple's research show almost zero emergent behaviour, so I'm inclined to think most of it was already in the weights.

It doesn't take away the usefulness, it just defined the boundary. We can't expect "original research" then because it actually can't reason about concepts that are too far from whats already in the discourse. The discourse is big so we don't notice.


You do realize that regardless of what was in the training data, the final solution included insights no human before had known, right? I share the same concerns regarding academic integrity but it would take a lot of motivated thinking to conclude that what the AI system did was not significant.

As far as I can tell (and my research was on the simulation side of Navier Stokes) the key AI output was a specific counter-example solution, generated with a method suspiciously close to that developed by the research duo involved in the controversy, a method that was discussed with Codex. So to me that insight is as insightful as the next undiscovered prime.

> People seem to be talking about anything except the actual results with this particular announcement.

To be fair, most people have a fairly good handle on "Does opting out my prompts from training runs actually work?", but not on Navier-Stokes. They discuss what more immediately affects them.


Additionally, I'm no physicist but I suspect the possibility of singularities in NS equations is probably one of those 'true but not meaningful' facts. If it took our brightest minds 175 years to craft such a scenario, how relevant can it be in practice? Especially when turbulence exists. Maybe I'm wrong or it has some consequences for pure math though.

Aren't you doing exactly the same thing as people you are mentioning? Skipping "talking about actual results" to talking about general capabilities of this LLM and computers in general? because that's exactly what seems like 99% of all people had been doing lately - debating what computer programs can do and what they can't.

Give me a dictionary, a computer, and infinite time, and I'll generate all possible English texts: Shakespeare, works regarded as surpassing Shakespeare, new holy books, math proofs never even imagined... none of which is either "creative" or "solving" anything. If I optimize my generation algorithm so that I'm not slavishly trying all possible combinations of words, it doesn't move me any closer to being creative, or solving anything.

The real casualty here may be our belief that humans are doing something more than some super-optimized version of what LLMs are doing. That doesn't elevate LLMs, it just makes us much less special.


>Its still astonishing that any sort of generalized computer program can solve a problem of this magnitude, and we have witnessed it happening in real time.

I think about this a lot. I'll have to explain to my kids some day that there was long period of time where you couldn't just talk to a computer and have it talk back to you, and that communicating with one required special skills that took years of study to master. It's going to be completely impossible for them to even remotely understand what that was like. Sort of like the pre-electricity days for us, but even more-so.


You're assuming you'll be the one doing the explaining :-)

It might also be that they won't even ask or wonder, similar to how most don't really do with pre-machining skills.

Or it could be like our "How did they build the Great Pyramid?!"


I mean, I have a bachelor's in math and I don't imagine I could begin to understand either the human or LLM proofs without a massive investment of time and effort.

But can you verify them? With less effort?

I'm not sure what you mean by "verify" here. I could run the lean verifier as could anyone else. Maybe I could write my own proof checker and do a purely mechanical translation into my own thing, though I don't think that's less effort.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: