Skip to content
Apixo
Blog
news· 4 min read· via TechCrunch AI

OpenAI's Math Proofs Fall Short of Mathematicians' Standards

OpenAI's release of hundreds of mathematical solutions has drawn criticism from researchers who warn of translation errors between natural language and code.

OpenAI's Math Proofs Fall Short of Mathematicians' Standards

OpenAI’s recent release of hundreds of proposed solutions to some of the world’s most challenging mathematics problems has sparked a critical response from the academic community. While the AI research lab sought to avoid previous controversies by consulting with elite mathematicians, experts argue that the company's release failed to meet the rigorous standards necessary for genuine scientific advancement. The core of the issue lies in the gap between automated outputs and actual human comprehension.

To guide AI laboratories, the Advisory Group on Mathematics and Artificial Intelligence (AGMAI)—a nine-member panel of prominent global researchers hosted by Princeton University’s Institute for Advanced Study—issued guidelines in late September. Despite this framework, OpenAI's latest release did not fully align with the group's recommendations. For instance, AGMAI’s primary request was for labs to stop testing advanced mathematical problems on proprietary, closed-source models. However, OpenAI explicitly evaluated its proprietary models using open math research problems.

Falling Short of Academic Guidelines

The discrepancies between the mathematical community's standards and OpenAI's release go beyond the use of proprietary systems. Although OpenAI followed some recommendations, such as publishing results quickly and providing details on how its models reached conclusions, it omitted critical data for the vast majority of its work. Out of 719 manuscripts released, only 10 included the models' chain-of-thought processes. Furthermore, 42% of the proofs remained unformalized, despite AGMAI’s recommendation that proofs not easily understood by humans should undergo formalization.

The group also requested that OpenAI include machine-readable metadata to connect natural language proofs with their formal counterparts. OpenAI did not provide this metadata. According to Harvard University mathematics professor Melanie Wood, the automated generation of these proofs does not equate to comprehension. "There is not human understanding of them at the point of release, and now the work begins," Wood noted, highlighting that human mathematicians must still do the heavy lifting to make these solutions useful.

The Translation Gap in AI Proofs

A significant technical concern involves how AI models translate natural language reasoning into formal code. When solving math problems, frontier models typically write a natural language explanation before translating it into Lean, a programming language designed to verify mathematical accuracy by compiling the proof as code.

However, a recent paper by researchers at the University of Cambridge and King’s College London revealed that this translation process is prone to errors. The researchers identified at least two distinct discrepancies between the natural language proof and the Lean code in OpenAI's solution to a problem derived from the Navier-Stokes equations, which describe fluid dynamics.

These translation errors mean the code and the natural language explanation do not fully align. The authors of the paper concluded that because of these mistranslations, natural language and autoformalized Lean proofs from OpenAI and other sources should not be trusted without the same level of peer review and scrutiny applied to human-generated proofs.

Prominent mathematician Terence Tao also expressed concern over this hands-off approach. He noted on social media that problems are being solved autonomously by AI prompters who lack interest in the broader mathematical field once a target is hit. These prompters often do not understand the AI's output well enough to answer questions, deliver presentations, or engage with other researchers.

What it means for developers

For software engineers and developers working on AI integration, these developments highlight the limitations of current LLM reasoning. While frontier models are increasingly capable of generating complex logic, the transition from natural language reasoning to formal, executable code remains highly vulnerable to translation errors. This is particularly critical for developers building automated verification systems, code generation tools, or mathematical modeling software.

Relying entirely on an AI's self-formalized code without human oversight or external peer review processes poses a high risk of introducing subtle, hard-to-detect bugs. Developers must design robust validation pipelines to verify that the generated code accurately reflects the intended natural language logic.

To explore how different LLMs handle complex reasoning and translation tasks, developers can try top AI models cheaply through one API at https://apixoai.online. Comparing the outputs of various proprietary models can help teams identify which architectures are best suited for structured logic and code generation while minimizing translation errors. Ultimately, human verification remains irreplaceable when deploying AI-generated logic in production environments.


Source: OpenAI’s math solutions aren’t meeting the field’s standards yet — TechCrunch AI. Written by the Apixo team from that report.

#ai-news#openai#artificial-intelligence#mathematics#llm#software-development
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading