[First key point: OpenAI published 722 math manuscripts from an unreleased model, but only 162 have been formally verified.][Second key point: MIT mathematician Andrew Sutherland calls the one-prompt, single-agent claim “unverified” and says the model must be released for replication.][Third key point: The Institute for Advanced Study warns that AI-produced math can be unverifiable, and OpenAI failed to disclose prompts as recommended.]
OpenAI published 722 math manuscripts on GitHub on Tuesday, all produced by an internal model the company has not released, claiming almost everything came from a single prompt handed to a single AI agent. The papers are grouped into 372 “families” of related results, bundling main theorems with companion arguments, consequences, or alternative proofs.
However, only 162 of the 722 papers include a computer-checked main result using Lean, software that mechanically verifies every logical step. OpenAI itself warns that “some of the unformalized results could have issues,” and a passing Lean check only confirms the proof follows from its stated premise, not whether the result is new or important.
Consequently, mathematicians are raising concerns. MIT’s Andrew Sutherland told Scientific American that until the model is released and results are replicated, claims about “one-shotting problems with a single agent” remain unverified. The Institute for Advanced Study in Princeton, New Jersey, issued a statement noting that AI can now output mathematical arguments “without the human who prompted it being able to understand the arguments, verify them, or take responsibility for them.”
Meanwhile, some researchers are excited. Professor Abhishek Saha called it “a very big day for mathematics,” though he noted most problems fit categories of “exceptional advances within an existing program” rather than game-changing breakthroughs. The release also falls short of advisory recommendations from the Institute for Advanced Study, which suggested disclosing the model name, prompts, chain of thought, time taken, and compute cost for every result.
OpenAI published average compute figures and reasoning summaries for only 10 results, omitting the prompts entirely. Anthropic took a different approach last month with its Lean-checked Fermat’s Last Theorem proof, posting all 13 million lines publicly on GitHub. OpenAI says it will add Lean formalizations as it obtains them, with 162 of 722 manuscripts currently having one.
✅ Follow BITNEWSBOT on Telegram, Facebook, LinkedIn, X.com, and Google News for instant updates.
