Eliza Ward: Hey. Good to be here.
Brian Reed: Yeah, likewise. You've been sitting on this one.
Eliza Ward: I have, yeah. Okay — so October 6, OpenAI drops 722 mathematical manuscripts on GitHub. Just... there. Public repo, Apache-2.0 license, organized into 372 result families. Claims solutions to, or — wait, the exact phrase is — 'significant progress on hundreds of long-standing open mathematics problems.'
Brian Reed: Hang on. None of them verified?
Eliza Ward: Not one. Not independently. Not on day of release.
Brian Reed: So the claims are — I mean, they're big claims sitting on a GitHub repo with no external mathematician having checked them yet. And Sam Altman is already calling it what?
Eliza Ward: 'A new era of discovery.' Same day. Before any independent review.
Brian Reed: That's a lot of weight to put on something nobody's had time to check.
Eliza Ward: And here's what the headline actually hides — you can read all 722 of those manuscripts and you still cannot reproduce a single result. Because the model that generated them? Unnamed. Not released. Locked.
Brian Reed: Wait — not named at all?
Eliza Ward: Not named. Think of it like — a dish arrives at your table. It's on the plate, you can taste it, but the kitchen is locked and the recipe doesn't exist. You can't check whether the cooking actually worked.
Brian Reed: Right — and that's the thing that makes this structurally different from a preprint. A preprint, I mean, it might be wrong, but you can follow the human mathematician's steps. Here the steps came from something OpenAI won't even name. So verification falls entirely on the reader, with no way to actually... retrace the path.
Eliza Ward: Lean, though. Isn't that the fix? Machine-checked proofs — isn't that exactly what Lean is for?
Brian Reed: That's what I thought, but — no, actually it doesn't cover it. Lean formalizations are present in selected manuscripts, not across all 722. And OpenAI's own formalization manifest — their own document — records partial progress with unchecked review status. They're flagging it themselves.
Eliza Ward: Unchecked. In their own manifest.
Brian Reed: Right. So the 722 number that's in every headline — some subset of those have partial Lean coverage, some have none, and OpenAI evaluated the underlying model on roughly 4,000 open problems to produce that pool. We don't know which of the 722 have real formal backing and which are... just text in a repo.
Eliza Ward: Which is the gap. Not 'math is hard to verify' — the specific gap is: anonymous model, partial Lean coverage, unchecked status in their own manifest. Those three together mean the openai/math repository is not self-contained evidence.
Brian Reed: But the take that's actually circulating — and I've seen it in a few places now — is that OpenAI did this responsibly. They consulted AGMAI at the Institute for Advanced Study. They disclosed compute cost. They included Lean proofs. That's the defense.
Eliza Ward: No. I don't buy it.
Brian Reed: Walk me through it.
Eliza Ward: Start with AGMAI. What did they actually do? OpenAI consulted them on best practices for releasing AI-generated results. That means the Advisory Group on Mathematics and Artificial Intelligence advised on the release format — how to package unverified work for public consumption. That is not the same as independently verifying whether a single one of those 722 manuscripts is mathematically correct. Those are completely different functions.
Brian Reed: So AGMAI's fingerprints are on the box, not the contents.
Eliza Ward: Exactly that. And then the compute disclosure — this is the one that actually bothers me. Three hours of ChatGPT Pro thinking per result, on average. Okay, that's a number. I can picture it. But there's no aggregate dollar cost, no token count, no hardware-hour total. They handed you one dimension of scale and stopped. The 10 reasoning summaries are the same thing — they're not proof verification, they're a narrative of how the model approached a problem. You read a summary and you think you've seen inside the process. You haven't.
Brian Reed: And then there's the Anthropic researcher calling this — wait, the actual quote — 'the most significant moment in mathematical history.' That landed before any external verification.
Eliza Ward: That's an opinion. Stated before verification. It's not an assessment — I mean, by what measure? More significant than Fermat's Last Theorem? Someone at a rival lab said it, and it became part of the framing. That's doing real work in how people receive this release.
Brian Reed: Which is the pattern across all three — AGMAI, the compute metric, the reasoning summaries. They perform transparency without actually delivering it. And what worries me is that if release-first becomes normal here, there's no institutional mechanism to clear the backlog — we'll get to exactly why that matters when we look at what this release changes about the timeline.
Eliza Ward: And that backlog is already three milestones deep. IMO gold medals — both DeepMind and OpenAI, summer 2025. Then September 8, 2026, OpenAI claims a solution to Navier–Stokes. One of the seven Clay Millennium Prize Problems. And then October 6 — 722 manuscripts. Each one arrived before the previous milestone had been verified.
Brian Reed: The Navier–Stokes one still isn't verified?
Eliza Ward: Still isn't. And the system that produced it — described only as significantly more capable than GPT-6 Astra. Also unnamed. Also unreleased.
Brian Reed: That actually breaks it open for me — OpenAI tested the underlying model on roughly 4,000 problems. Released 722. Which means... I mean, what happened to the other 3,300? Are those coming? Is there a second batch? A larger one? And the math community has no process that runs at the speed of a GitHub push.
Eliza Ward: Wait — that's the number stuck with me. 4,000 problems evaluated, 722 released. That gap isn't random.
Brian Reed: Right, and let me make this concrete. Say you're a PhD student, it's a Wednesday night, your advisor just forwarded you the openai/math repository. There's a manuscript — call it Manuscript 47 — that overlaps your dissertation chapter. Do you cite it? Do you spend the next three weeks re-verifying it from scratch? Because there's no institutional answer. No journal said it's peer reviewed. AGMAI advised on the packaging, not the contents. And Lean coverage is partial — OpenAI's own manifest says so.
Eliza Ward: And if she builds on it and a flaw surfaces two years later — you can't un-cite something.
Brian Reed: That's exactly it. Anthropic has results in this space too, which means this isn't one lab's problem — the production rate is competitive across multiple organizations simultaneously. The concrete thing to watch is whether any Clay Institute adjudicator formally engages with the Navier–Stokes claim before the next batch drops. Because if batch two lands first, the queue problem compounds.
Eliza Ward: That's the signal. Not 'will mathematicians accept this' — it's whether any formal verification catches up before the next release resets the clock entirely.
Brian Reed: And I don't think mathematics as an institution has decided yet whether it's going to hold the line on that. Standard publishing norms — independent verification before a claim is treated as established — those norms are what the October 6 release explicitly bypasses. But does the field push back? Does it adapt some faster track? Or does it just... absorb the backlog because the results keep coming faster than any review process can run?
Eliza Ward: I mean — that last one is the one that actually worries me. Not 'absorb and catch up.' Absorb and never catch up. Because if release-first becomes the norm here, you end up with mathematics as a repository of unvalidated claims waiting on human review that may not arrive at anywhere near the speed AI can produce.
Brian Reed: Which is — wait, that's a genuinely different thing from 'peer review is slow.' This is scale outrunning the institution entirely.
Eliza Ward: Right. And what I don't know — what I actually can't answer — is whether the Clay Mathematics Institute or AGMAI or anyone with institutional standing will say, formally: Navier–Stokes is unverified, these 722 manuscripts are unverified, and we are not treating speed as a substitute for certainty. That's the question. Not whether the model is impressive. It probably is. It's whether the field holds its standard or quietly lets it slip.
Brian Reed: We genuinely don't have that answer yet.