The average result in the 722-manuscript release took roughly the compute of three hours of ChatGPT Pro thinking, a small fraction of what the company says it spent on its Navier–Stokes claim.
OpenAI has disclosed how much compute went into its latest batch of mathematical results, and the number is small. The average result used the equivalent of roughly three hours of ChatGPT Pro thinking, the company says.
“To promote scientific transparency and openness, we are also publishing additional details about how we obtained the results in the repository,” OpenAI said in its announcement. These include 10 summaries of the model’s reasoning, estimates of compute spent in terms of Pro usage on ChatGPT, and statistics about the number of attempted problems. The company also states that “the average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.”

What the numbers say
The release is a catalog of 722 manuscripts grouped into 372 result families, covering topics from number theory to partial differential equations. According to the repository’s readme, the model was posed approximately 4,000 problems over the course of the evaluation. OpenAI says it expanded these evaluations after its existing math evaluations saturated, then aggregated the outputs into families and manuscripts and required “an appropriate level of significance,” which produced the final catalog.
Most results, OpenAI says, came from one fixed procedure with an unreleased internal model. The readme names two exceptions to that procedure: the work on a zero-free region for the Riemann zeta function, and the proof of the Hodge conjecture for CM abelian varieties. It doesn’t say what was different about how those two were produced, or how much compute they used. The writeup of the Re(s) > 11/12 zero-free region was also edited by humans for readability.
Two things the figure doesn’t tell you are worth keeping in mind. First, it is an average, and the repository’s own wording is “the average result”, so some results presumably used far more and some far less. Second, it covers results that made the cut. Roughly 4,000 problems were attempted and 372 families emerged, so the cost of the problems that didn’t produce a result isn’t captured in a per-result figure. OpenAI says it is publishing statistics on the number of attempted problems, which is the number needed to work out the cost of a success, but I haven’t gone through those in detail.
How it compares
The three-hour figure is a deliberate contrast with OpenAI’s biggest recent claim. When the company said an internal model had solved the Navier–Stokes Millennium Prize problem, reports put the effort at roughly 10,000 agents working in parallel for about 88 hours, at a cost reported to run into the millions of dollars and consuming around 130 billion output tokens. A single ChatGPT Pro session is a different order of magnitude.
It is also in the same range as other recent claims about quick AI proofs. OpenAI said GPT 5.6 Sol produced a proof of the 50-year-old cycle double cover conjecture using 64 subagents in one hour, and DeepMind’s AlphaProof Nexus agent solved nine open Erdős problems at a cost of a few hundred dollars each.
Why it matters
The per-result cost is arguably as important to the industry as the results. If a typical research-level result takes hours of a consumer-tier product’s reasoning, rather than weeks of a cluster, then the supply of such results is mostly a function of how many problems get posed and who checks the answers. That also helps explain why OpenAI’s tally of solved problems has climbed so fast: an informal tracker counted 234 problems solved with OpenAI’s models as of early August, ahead of Anthropic’s 57, and the company has since said its models have resolved more than 100 long-standing problems beyond Navier–Stokes.
It also sharpens a concern raised by mathematicians. If results are this cheap to generate, verification becomes the bottleneck. On September 11, 25 Fields Medal winners, including Terence Tao, signed a declaration warning that rapid AI proofs are being announced in a rush, without proper verification, and Tao has said such proofs could be a net negative for mathematics if they arrive as opaque black boxes. OpenAI’s own readme says not every manuscript has a Lean formalization and that some unformalized results “could have issues.” None of the results has been confirmed by outside mathematicians, and the three-hour figure says how little compute produced the claims, not how likely they are to be right.
What to watch
The numbers that matter next are the ones that aren’t in the headline: how many of the roughly 4,000 problems were attempted more than once, how the cost of the two exception results compares with the average, and how many of the 722 manuscripts survive scrutiny. The reasoning summaries for 10 of the families, including the quasi-Riemann hypothesis, the Unique Games Conjecture and the free group factors, give outside mathematicians a first look at how the model got there.