I think too much credit in the AI math discussion is being given to LLMs rather than to Lean, which an incredibly well designed language around which Mathlib coalesced as a side-effect of it's capability
I doubt of this progress could have been made without the specific combination of Lean + Mathlib. Automatic Theorem Provers are not a new idea, but we don't see these breakthroughs happening with any other stack
Indeed, I don't think anyone should be led to believe that AI can do math on its own without _some_ kind of oracle providing extremely rigorous feedback.
Not sure why you say that. All the major AI-generated proofs so far were done by AI without using lean, and then later validated by AI formalizing with lean in a separate process. Do you mean we would not have enough rigorous training data perhaps?
because LLMS are like drunk geniuses, without the rigor of lean they would drift into nonsense. Similar in chess they have 1700 ELO now but still some percent of their moves are illegal.
> I think too much credit in the AI math discussion is being given to LLMs rather than to Lean,
For sure math and software people know this, as it's kind of hard to ignore.
Developers who are totally disinterested in nerdy stuff like theorem proving will very quickly hit a wall without letting their models indulge in test suites (regardless of whether they ever even glance at the tests themselves). Physics nerds get it, and if they are on more on the experimental side then they are working to give it some kind of sandbox, all the other science wonks are getting it or will get it as they work with their more computationally inclined colleagues, same for general engineering.
There's a real danger here though in that there is no reason the general public will ever become very much acquainted with the verification/validation steps that all these systems actually need to be reliable. They just hear about this or that nearly impossible thing finally being accomplished as if by magic, and they won't do the work to think about what's in scope, out of scope, or how anything can be checked iteratively.
I am not really sure. The models can be bootstrapped to become much better at both natural language generation and verification for sure and my hypothesis is that the current days models would have been just as good at math without Lean.
Relatedly, I think where I land is that the (current) challenge for humans is creating such oracles. AIs have patience and compute we can only dream of—I think someone estimated 100,000 person-years were spent on Navier-Stokes—but they still need a “reality” to measure against.
I doubt of this progress could have been made without the specific combination of Lean + Mathlib. Automatic Theorem Provers are not a new idea, but we don't see these breakthroughs happening with any other stack