OpenAI’s Math Solutions Aren’t Meeting the Field’s Standards Yet
The frontier lab released hundreds of claimed solutions to hard math problems this week, but mathematicians say the work falls short on the very principles OpenAI claimed to be following.
The Promise and the Problem
When OpenAI unveiled hundreds of purported solutions to some of mathematics’ most challenging open problems this week, the company framed the release as a careful, responsible step forward. Having faced controversy before over claims that its models had solved long standing problems, OpenAI said it had consulted an advisory group of elite mathematicians this time around.
But according to those very mathematicians, the release fell short, particularly on the question of whether humans actually understand the results.
What the Advisory Group Asked For
The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by the Institute for Advanced Study and composed of nine prominent researchers from institutions worldwide, released guidelines for frontier labs at the end of September. Among its recommendations:
- Stop testing advanced mathematical problems on proprietary models
- Release results as soon as possible, along with information about how models reached their conclusions
- Formalize proofs that people don’t understand
- Include machine readable metadata correlating natural language and formal artifacts
- Take responsibility for ensuring human understanding follows
OpenAI’s release explicitly states that it is evaluating its proprietary models using open research problems in mathematics, directly contradicting the group’s first request.
The Numbers Tell a Story
OpenAI clearly followed some of AGMAI’s principles. But the gaps are striking:
| AGMAI Recommendation | OpenAI’s Compliance |
|---|---|
| Release results promptly | Followed |
| Include reasoning information | Only 10 of 719 manuscripts included chain of thought |
| Formalize unclear proofs | 58% remain unformalized |
| Include metadata linking natural language to formal proofs | Not done |
The “Lost in Translation” Problem
A new paper from mathematicians at the University of Cambridge and King’s College London highlights a critical flaw in how AI models approach mathematical proof. The process typically works in two stages:
- The model generates a natural language explanation of the proof
- The model attempts to express that result in Lean, a programming language that verifies accuracy by compiling the proof as code
But the translation between these two stages can introduce errors. The paper documents at least two discrepancies between the natural language proof and the Lean code behind OpenAI’s solution to a problem derived from the Navier Stokes equations, the notoriously difficult equations describing fluid behavior.
These discrepancies don’t necessarily disprove either solution. But they raise a troubling question: Can we trust models to formalize their own solutions without human oversight?
As the paper’s authors conclude:
“Because of the phenomenon of mistranslations… the NL proof by OpenAI and other autoformalised Lean proofs should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to.”
The Human Understanding Gap
Perhaps the most fundamental concern raised by mathematicians is about understanding, not just correctness.
When human mathematicians discover new results, they take responsibility for them. They engage with the broader community through papers, talks, and seminars. This process:
- Increases understanding of the solutions
- Reveals strategies applicable to other problems
- Allows new knowledge to be applied in practical fields
When an AI model spits out a solution to a hard problem, none of this happens automatically.
“There is not human understanding of them at the point of release, and now the work begins.”
— Melanie Wood, Harvard University mathematics professor
Terence Tao, one of the most prominent mathematicians to criticize OpenAI’s approach, put it bluntly on social media:
“Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is ‘solved’, and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field.”
What Comes Next?
AGMAI has suggested that OpenAI should help fund the work of human mathematicians who will be required to make the lab’s solutions meaningful. So far, there’s no indication that has happened.
The advisory group, for its part, offered a measured statement on the latest proofs:
“It is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully.”
That assessment, based on the evidence so far, appears to be: not very well.
Conclusion
OpenAI’s latest math release represents a genuine technical achievement. Solving hundreds of open problems is no small feat. But solving problems and advancing mathematics are not the same thing.
Mathematics is not merely a collection of correct answers. It is a living discipline built on understanding, communication, and communal verification. When AI generates proofs that no human fully comprehends, and when the formalization process itself introduces potential errors, the field is left with results it cannot yet trust.
The path forward, as AGMAI and the mathematicians quoted here suggest, requires more than just better models. It requires human involvement at every stage, from problem selection to proof verification to the slow, essential work of building understanding. Until then, OpenAI’s math solutions will remain what they currently are: impressive outputs that have not yet met the standards of the field they claim to advance.
TechTrib.com is a leading technology news platform providing comprehensive coverage and analysis of tech news, cybersecurity, artificial intelligence, and emerging technology. Visit techtrib.com.
Contact Information: Email: [email protected] or for adverts placement [email protected]