Onpode
Cover art for OpenAI just published 722 AI-written math papers solving hundreds of open problems

OpenAI just published 722 AI-written math papers solving hundreds of open problems

October 7, 2026 · 10 min

Eliza Ward & Brian Reed

On October 6, OpenAI published 722 AI-written math manuscripts on GitHub — organized into 372 result families and claiming progress on hundreds of open problems — but not one manuscript had been independently verified at release, and the model that generated them was neither named nor released.

On October 6, 2026, OpenAI announced the public release of 722 mathematical manuscripts produced by an unnamed, unreleased internal frontier model.

0:009:45
Get the next episode on Technology →

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Technology →

About this episode

On October 6, 2026, OpenAI published 722 AI-generated mathematical manuscripts to a public GitHub repository, organized into 372 result families, claiming significant progress on hundreds of long-standing open problems. The model that produced them was unnamed and unreleased. Not one manuscript had been independently verified at the time of publication. This episode works through what that actually means — and why the framing around the release matters as much as the release itself. The AGMAI advisory group at the Institute for Advanced Study is cited as evidence of responsible disclosure, but their role was advising on how to package unverified work, not checking whether any of it is correct. Lean proof formalizations appear in selected manuscripts, not all 722, and OpenAI's own manifest records unchecked status. The compute disclosure gives you one number — roughly three hours of ChatGPT Pro thinking per result — but no aggregate cost, no token count, no hardware total. Underneath all of that is a harder structural problem: OpenAI evaluated the underlying model on approximately 4,000 open problems and released 722. The September 2026 Navier–Stokes claim — one of the seven Clay Millennium Prize Problems — still hasn't been formally verified. Each new release arrives before the previous milestone has cleared. The episode asks whether any institution with formal standing will hold the line on verification before the next batch resets the clock entirely — and whether 'absorb and never catch up' is already the default path.

Frequently asked

What did OpenAI release in its 722 math manuscripts?

OpenAI published 722 AI-generated mathematical manuscripts on GitHub, organized into 372 result families, claiming significant progress on hundreds of long-standing open problems. The repository is licensed under Apache-2.0. No external mathematician had independently verified any of the manuscripts at the time of release.

Were OpenAI's 722 math manuscripts independently verified?

None of OpenAI's 722 math manuscripts were independently verified at release. Lean formal proof coverage is partial, not universal across all 722 papers, and OpenAI's own formalization manifest records unchecked review status on multiple entries — meaning even OpenAI flagged incomplete verification in its own documentation.

What role did AGMAI play in OpenAI's math manuscript release?

AGMAI — the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study — advised OpenAI on best practices for releasing AI-generated results. That means AGMAI guided the release format and packaging, not the mathematical correctness of the manuscripts. AGMAI did not independently verify any of the 722 claims.

Has OpenAI's claimed Navier–Stokes solution been verified?

OpenAI's claimed solution to the Navier–Stokes Millennium Prize Problem, announced September 8, 2026, remained unverified as of the October 6 manuscript release. The model that produced it was described only as 'significantly more capable than GPT-6 Astra' and was neither named nor made publicly available for independent replication.

Why can't researchers reproduce results from OpenAI's math manuscripts?

Researchers cannot reproduce OpenAI's math manuscript results because the AI model that generated all 722 papers is unnamed and unreleased. Without access to the model or its methodology, readers can read the outputs but cannot retrace the generative steps — a structural difference from conventional mathematical preprints.

Grounded in 10 sources
Learning Strategies To Break Judges ↗ · arxiv.org
OpenAI announces 722 mathematical discoveries in one go | New Scientist ↗ · newscientist.com
OpenAI drops another batch of mathematical breakthroughs | The Verge ↗ · theverge.com
OpenAI's largest math release tackles 4,000 problems with Lean proofs ↗ · interestingengineering.com
OpenAI’s 722 Math Manuscripts: The Results, Proofs, Compute and Costs ↗ · kingy.ai
[AINews] Quasi-Riemann-Hypothesis: OpenAI publishes 722 math papers solving 90 of the top 500 open math problems; “the most significant moment” in >100 years of mathematics ↗ · latent.space
Sharing AI progress in mathematics ↗ · openai.com
OpenAI Math Advisory Group: 100+ Claims Unverified ↗ · shattered.io
OpenAI Releases 722 Math Manuscripts From Hidden Model ↗ · shattered.io
OpenAI Releases 722 Math Manuscripts From an Unreleased AI Model – Unite.AI ↗ · unite.ai
Read transcript

Eliza Ward: Hey. Good to be here.

Brian Reed: Yeah, likewise. You've been sitting on this one.

Eliza Ward: I have, yeah. Okay — so October 6, OpenAI drops 722 mathematical manuscripts on GitHub. Just... there. Public repo, Apache-2.0 license, organized into 372 result families. Claims solutions to, or — wait, the exact phrase is — 'significant progress on hundreds of long-standing open mathematics problems.'

Brian Reed: Hang on. None of them verified?

Eliza Ward: Not one. Not independently. Not on day of release.

Brian Reed: So the claims are — I mean, they're big claims sitting on a GitHub repo with no external mathematician having checked them yet. And Sam Altman is already calling it what?

Eliza Ward: 'A new era of discovery.' Same day. Before any independent review.

Brian Reed: That's a lot of weight to put on something nobody's had time to check.

Eliza Ward: And here's what the headline actually hides — you can read all 722 of those manuscripts and you still cannot reproduce a single result. Because the model that generated them? Unnamed. Not released. Locked.

Brian Reed: Wait — not named at all?

Eliza Ward: Not named. Think of it like — a dish arrives at your table. It's on the plate, you can taste it, but the kitchen is locked and the recipe doesn't exist. You can't check whether the cooking actually worked.

Brian Reed: Right — and that's the thing that makes this structurally different from a preprint. A preprint, I mean, it might be wrong, but you can follow the human mathematician's steps. Here the steps came from something OpenAI won't even name. So verification falls entirely on the reader, with no way to actually... retrace the path.

Eliza Ward: Lean, though. Isn't that the fix? Machine-checked proofs — isn't that exactly what Lean is for?

Brian Reed: That's what I thought, but — no, actually it doesn't cover it. Lean formalizations are present in selected manuscripts, not across all 722. And OpenAI's own formalization manifest — their own document — records partial progress with unchecked review status. They're flagging it themselves.

Eliza Ward: Unchecked. In their own manifest.

Brian Reed: Right. So the 722 number that's in every headline — some subset of those have partial Lean coverage, some have none, and OpenAI evaluated the underlying model on roughly 4,000 open problems to produce that pool. We don't know which of the 722 have real formal backing and which are... just text in a repo.

Eliza Ward: Which is the gap. Not 'math is hard to verify' — the specific gap is: anonymous model, partial Lean coverage, unchecked status in their own manifest. Those three together mean the openai/math repository is not self-contained evidence.

Brian Reed: But the take that's actually circulating — and I've seen it in a few places now — is that OpenAI did this responsibly. They consulted AGMAI at the Institute for Advanced Study. They disclosed compute cost. They included Lean proofs. That's the defense.

Eliza Ward: No. I don't buy it.

Brian Reed: Walk me through it.

Eliza Ward: Start with AGMAI. What did they actually do? OpenAI consulted them on best practices for releasing AI-generated results. That means the Advisory Group on Mathematics and Artificial Intelligence advised on the release format — how to package unverified work for public consumption. That is not the same as independently verifying whether a single one of those 722 manuscripts is mathematically correct. Those are completely different functions.

Brian Reed: So AGMAI's fingerprints are on the box, not the contents.

Eliza Ward: Exactly that. And then the compute disclosure — this is the one that actually bothers me. Three hours of ChatGPT Pro thinking per result, on average. Okay, that's a number. I can picture it. But there's no aggregate dollar cost, no token count, no hardware-hour total. They handed you one dimension of scale and stopped. The 10 reasoning summaries are the same thing — they're not proof verification, they're a narrative of how the model approached a problem. You read a summary and you think you've seen inside the process. You haven't.

Brian Reed: And then there's the Anthropic researcher calling this — wait, the actual quote — 'the most significant moment in mathematical history.' That landed before any external verification.

Eliza Ward: That's an opinion. Stated before verification. It's not an assessment — I mean, by what measure? More significant than Fermat's Last Theorem? Someone at a rival lab said it, and it became part of the framing. That's doing real work in how people receive this release.

Brian Reed: Which is the pattern across all three — AGMAI, the compute metric, the reasoning summaries. They perform transparency without actually delivering it. And what worries me is that if release-first becomes normal here, there's no institutional mechanism to clear the backlog — we'll get to exactly why that matters when we look at what this release changes about the timeline.

Eliza Ward: And that backlog is already three milestones deep. IMO gold medals — both DeepMind and OpenAI, summer 2025. Then September 8, 2026, OpenAI claims a solution to Navier–Stokes. One of the seven Clay Millennium Prize Problems. And then October 6 — 722 manuscripts. Each one arrived before the previous milestone had been verified.

Brian Reed: The Navier–Stokes one still isn't verified?

Eliza Ward: Still isn't. And the system that produced it — described only as significantly more capable than GPT-6 Astra. Also unnamed. Also unreleased.

Brian Reed: That actually breaks it open for me — OpenAI tested the underlying model on roughly 4,000 problems. Released 722. Which means... I mean, what happened to the other 3,300? Are those coming? Is there a second batch? A larger one? And the math community has no process that runs at the speed of a GitHub push.

Eliza Ward: Wait — that's the number stuck with me. 4,000 problems evaluated, 722 released. That gap isn't random.

Brian Reed: Right, and let me make this concrete. Say you're a PhD student, it's a Wednesday night, your advisor just forwarded you the openai/math repository. There's a manuscript — call it Manuscript 47 — that overlaps your dissertation chapter. Do you cite it? Do you spend the next three weeks re-verifying it from scratch? Because there's no institutional answer. No journal said it's peer reviewed. AGMAI advised on the packaging, not the contents. And Lean coverage is partial — OpenAI's own manifest says so.

Eliza Ward: And if she builds on it and a flaw surfaces two years later — you can't un-cite something.

Brian Reed: That's exactly it. Anthropic has results in this space too, which means this isn't one lab's problem — the production rate is competitive across multiple organizations simultaneously. The concrete thing to watch is whether any Clay Institute adjudicator formally engages with the Navier–Stokes claim before the next batch drops. Because if batch two lands first, the queue problem compounds.

Eliza Ward: That's the signal. Not 'will mathematicians accept this' — it's whether any formal verification catches up before the next release resets the clock entirely.

Brian Reed: And I don't think mathematics as an institution has decided yet whether it's going to hold the line on that. Standard publishing norms — independent verification before a claim is treated as established — those norms are what the October 6 release explicitly bypasses. But does the field push back? Does it adapt some faster track? Or does it just... absorb the backlog because the results keep coming faster than any review process can run?

Eliza Ward: I mean — that last one is the one that actually worries me. Not 'absorb and catch up.' Absorb and never catch up. Because if release-first becomes the norm here, you end up with mathematics as a repository of unvalidated claims waiting on human review that may not arrive at anywhere near the speed AI can produce.

Brian Reed: Which is — wait, that's a genuinely different thing from 'peer review is slow.' This is scale outrunning the institution entirely.

Eliza Ward: Right. And what I don't know — what I actually can't answer — is whether the Clay Mathematics Institute or AGMAI or anyone with institutional standing will say, formally: Navier–Stokes is unverified, these 722 manuscripts are unverified, and we are not treating speed as a substitute for certainty. That's the question. Not whether the model is impressive. It probably is. It's whether the field holds its standard or quietly lets it slip.

Brian Reed: We genuinely don't have that answer yet.

OpenAI just published 722 AI-written math papers solving hundreds of open problems · Onpode