OpenAI published a collection of mathematical results from an internal model on October 6. Its public repository contains 722 manuscripts arranged into 372 families of related work. The repository says the model was presented with about 4,000 problems during the evaluation.
OpenAI describes the system as an unreleased internal model. The repository says most results followed one procedure, with some exceptions. It reports an average of about three hours of ChatGPT Pro thinking compute per result. That figure describes the company's process. It is not a benchmark against human researchers.
1,000+ Proven ChatGPT Prompts That Help You Work 10X Faster
ChatGPT is insanely powerful.
But most people waste 90% of its potential by using it like Google.
These 1,000+ proven ChatGPT prompts fix that and help you work 10X faster.
Sign up for Superhuman AI and get:
1,000+ ready-to-use prompts to solve problems in minutes instead of hours—tested & used by 1M+ professionals
Superhuman AI newsletter (3 min daily) so you keep learning new AI tools & tutorials to stay ahead in your career—the prompts are just the beginning
The materials include manuscript files, reasoning summaries, and Lean formalizations for some results. Lean is a proof assistant that checks proofs expressed in its formal language. OpenAI says the collection contains outputs at different verification stages. Many manuscripts have formalizations, but not all.
The repository also warns that some results without formal proofs could contain problems. It says corrections will be recorded in later versions, while earlier releases remain accessible. A formalized proof can be checked within the stated formal system, but the repository does not present every manuscript as mechanically verified.
That distinction changes how the headline number should be read. The collection includes related papers, companion arguments, consequences, and alternative proofs. Therefore 722 manuscripts are not necessarily 722 independent discoveries. The repository groups related work into families for that reason.
OpenAI says it consulted an independent advisory group at the Institute for Advanced Study while preparing its release practices. The group responded on October 6. It said its advisory role should not be read as endorsing the results or the process used to produce them. It said mathematicians must assess the work themselves.
10 AI Stocks Investors May Regret Ignoring
AI has already produced some of the stock market’s biggest winners, but the opportunity may be far from over.
This free report reveals 10 AI stocks positioned for the next wave of investment and explains how each business makes money, what could drive growth and which risks investors should understand.
The group called public access a first step rather than the completion of review. It said mathematicians need the freedom and resources to develop their own questions, not only evaluate problems selected by AI laboratories. The group also said the community must assess whether its release recommendations were followed. Those are views about research priorities and process, not verdicts on each proof.
OpenAI's repository provides material for inspection. It includes preprints, formal proof files, and a catalogue explaining the collection. Those artifacts make some parts of the work more checkable than a bare capability claim. They do not settle whether every result is correct, new, important, or well explained.
The distinction between formal correctness and mathematical contribution matters here. A proof checker evaluates a formal statement encoded for it. Human reviewers still need to compare that statement with the paper's argument and prior literature. They also need to assess the intended question. The advisory group leaves that work to the mathematical community.
The release also contains exceptions to the standard process. OpenAI says work on a region without zeros of the Riemann zeta function and a proof related to the Hodge Conjecture used different procedures. It notes that one writeup was edited by a human for readability. The repository preserves earlier versions so later corrections do not erase the release history.
These details argue against treating the entire collection as one uniform experiment. The number of manuscripts, amount of compute, and mix of verification stages describe a research release with varied evidence. They do not establish a general intelligence threshold or prove that all listed problems have been independently settled.
The supported conclusion is narrow: an internal model produced substantial mathematical work, and OpenAI has exposed artifacts for examination. Independent checking, not the upload's size, will determine what the results establish. Their significance remains open.
The agentic era needs a different CRM. That’s Attio.
Parallel, Turbopuffer, and Wordsmith run their entire GTM motion on Attio, with agents that chase every buying signal, build pipeline, and move deals forward, 24/7.





