Felix Ortiz: Hey, long week — are you still processing the thing I sent you, or have you already built a whole thesis about it?
Tess Hollis: Both, honestly. The July 2026 proof collision has been sitting in my head like a splinter.
Felix Ortiz: Okay good, because here is the part that I think people are glossing over — it's not that two teams got close to the same answer. Seyoon Ragavan at MIT and a cryptography group at UC Santa Barbara both used GPT-5.6 Sol Ultra, both proved the exact same quantum cryptography problem, and the papers landed three hours apart. The proofs were identical.
Tess Hollis: Identical like — same logical steps, same structure, same output?
Felix Ortiz: Same proof. Which means — and I keep turning this over — access to the model is now the discovery. That's it. You have GPT-5.6 Sol Ultra and the right prompt, you have the result.
Tess Hollis: So the question that had never needed asking — who discovered this — suddenly has no answer.
Felix Ortiz: Or the answer is just: whoever submitted first. Which is a terrifying thing to say about a scientific result.
Tess Hollis: And that's where I want to start — because I think this is the moment the credit system actually broke, not bent.
Felix Ortiz: But here's what I keep wanting to back up and say — the collision didn't come from nowhere. Like, that moment makes way more sense if you trace what actually happened in the two years before it.
Tess Hollis: The IMO run. Yeah. Walk it back.
Felix Ortiz: Okay — so imagine you hand a chess puzzle to a friend, a puzzle that stumped grandmasters for years. They solve it overnight. That's roughly AlphaProof and AlphaGeometry 2 at the 2024 IMO — four of six problems, silver-medal threshold. And I mean, that's real, that was exciting. But — wait, actually the part people glossed over — they needed human experts to convert the problems into machine-readable form first. Autoformalization wasn't solved. It was a mediated result.
Tess Hollis: And it took multiple days of computation. So not exactly overnight.
Felix Ortiz: Right — but then 2025, different friend, same puzzle, solved in an afternoon, and they didn't even read the instructions first. Gemini Deep Think, five of six, gold-medal threshold, natural language directly. No human translation layer at all.
Tess Hollis: That gap — mediated to direct — that's not incremental. That's a qualitative shift. The system stopped needing a human in the loop to even understand the question.
Felix Ortiz: And then — I mean, this is the part that kind of breaks the whole benchmark framing — AlphaEvolve moved beyond competition math entirely. Research-level open problems. So the IMO, which felt like this huge milestone in 2024, already feels almost quaint by mid-2026.
Tess Hollis: The benchmark saturated before anyone agreed what it was measuring.
Felix Ortiz: And that's — yeah, that's the actual new thing. Not that AI can do math. It's that the speed from 'silver medal with help' to 'gold medal alone' to 'beyond competition math entirely' was, what, roughly two years? Nobody knows what benchmark comes next, and history suggests AI will saturate it before the math community even votes on whether it counts.
Tess Hollis: But that's exactly where the circulating take breaks down — because a lot of people heard 'Lean verified it' and just... stopped. Like, proof assistant checked every step against formal logical rules, machine-verified certificate of correctness, case closed.
Felix Ortiz: Right — and I want to push on that, because I actually think that's — wait, no, let me back up — the argument isn't wrong, it's incomplete. Lean proves the logical chain works. That's real. But AlphaProof learns by reinforcement learning, reward signals for correct steps. It's optimizing toward outputs that pass formal checks. That tells us nothing about whether anything was understood.
Tess Hollis: Okay, but — what standard beyond logical correctness have mathematicians ever formally required? That's the proponent's actual point and it has teeth.
Felix Ortiz: Oh — wait, that's — yeah, that does land.
Tess Hollis: A proof is either right or it isn't. That's been the standard. So if you're saying Lean-verified proofs don't count, you're adding a requirement that wasn't on paper before.
Felix Ortiz: Okay, I'll concede that — but here's where I think the community is actually split, and I mean genuinely split, not just rhetorically. Some people are saying formal verification is the finish line. Others are saying it's the starting point — that validity and discovery are different questions entirely. A proof no human can follow might be logically sound but mathematically sterile, and those are not the same failure.
Tess Hollis: Pattern-matching versus insight. Yeah. And neither side has settled it.
Felix Ortiz: Nobody's settled it — and I don't think the research settles it cleanly either, I mean I genuinely don't know how you'd even measure 'understood.' But the downstream of that unresolved question is where it gets worse — the Leiden Declaration, the unit-distance problem, who actually controls which problems get worked on. That's coming and it's — yeah, it's a different kind of uncomfortable.
Tess Hollis: Verification answered the wrong question. And nobody wanted to say it out loud.
Felix Ortiz: Right — and the wrong question is now embedded in an institution, which is — okay, that's where the unit-distance thing lands so hard. OpenAI announces AI solved it. Not a math department. Not a journal. A tech company put out a press release saying geometry's unit-distance problem is done.
Tess Hollis: And the alarm wasn't about the math being wrong.
Felix Ortiz: No — it was about who made the announcement. Like, that's — wait, that's actually the whole thing. The result could be correct and the community is still right to be alarmed, because the pipeline just bypassed every gate they'd built.
Tess Hollis: Which is what the Leiden Declaration is actually about. Not 'AI proofs are fake.' The structural concern is that Google DeepMind and OpenAI own the most powerful tools, so they're increasingly deciding which problems get worked on, how results get disseminated. That's not philosophy — that's a power question.
Felix Ortiz: Okay but — is that turf war dressed in philosophical language, or is it a real structural problem? Because I genuinely can't tell how much to weight it.
Tess Hollis: Probably both — and the mixing is the real story. But here's what makes me take the structural side seriously: imagine a postdoc at a mid-tier university, November 2026, submitting to a journal. The flood of low-quality AI-assisted submissions — which the Declaration and Nature commentary both flagged as a specific documented harm — means reviewers are already overwhelmed, standards are already slipping. Her careful work goes into that same pile.
Felix Ortiz: And she either uses the same tool that created the flood, or she falls behind people who did.
Tess Hollis: That's the Tuesday-morning version of this. Not abstract. Who gets hired, what counts as a contribution — those answers are now downstream of which tools two companies decide to release.
Felix Ortiz: And the Leiden Declaration's fifteen co-signatories are trying to build governance before there's anything to govern with. I mean — I genuinely don't know if that's brave or just late.
Tess Hollis: Maybe that's the actual question the field hasn't answered yet. Not 'is the proof valid' — Lean can settle that. But whether mathematics is fundamentally about accumulating true statements, or about humans understanding why they're true. Because those are different projects. And if it's the first one, AI abundance is a golden age. If it's the second, it's — I don't know what it is.
Felix Ortiz: I don't think the field has ever had to answer that. Like, it never came up because the human understanding and the proof generation were always bundled together. You couldn't have one without the other. Now you can, and nobody agreed on which one was the point.
Tess Hollis: The Leiden Declaration is one attempt to force that choice. But the window to make it consciously — I genuinely don't know how open it still is.
Felix Ortiz: Yeah. I don't either, actually.
Tess Hollis: Good conversation to not resolve.