Onpode
Cover art for Two research teams just used AI to prove the same open cryptography problem — three hours apart

Two research teams just used AI to prove the same open cryptography problem — three hours apart

July 31, 2026 · 10 min

Archer Whitfield

On July 25, 2026, two separate research teams — Seyoon Ragavan at MIT and a UCLA/UC Santa Barbara group — posted proofs of the same open unclonable cryptography problem to arXiv within three hours of each other. Both had used GPT-5.6 Sol Ultra. That coincidence has fractured the scientific concept of independent discovery.

On July 25, 2026, MIT Ph.D. student Seyoon Ragavan and UC Santa Barbara professor Prabhanjan Ananth (along with doctoral student Yao-Ting Lin) and UCLA professor Amit Sahai each independently used GPT-5.6 Sol Ultra to solve the same open problem in quantum cryptography — a specialized area known as unclonable cryptography.

0:009:57
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

On July 25, 2026, two preprints appeared on arXiv within three hours of each other, both claiming to solve the same open problem in unclonable cryptography — the branch of quantum cryptography concerned with objects that cannot be copied by any process quantum mechanics allows. One came from a solo MIT PhD student. The other from a UCLA/UC Santa Barbara group. Both teams had encountered the problem at the same workshop at the Simons Institute for the Theory of Computing earlier that month, and both had used GPT-5.6 Sol Ultra to solve it. This episode works through why that three-hour gap is not evidence of independence — and why that distinction matters. Scientific priority, the entire system of credit and verification, rests on independence as a load-bearing assumption. When both paths to a result run through the same model, you're not looking at two paths. You're looking at two entry points to the same structure. Three days later, Anthropic disclosed that Claude Mythos Preview had found a structural flaw in HAWK, a NIST post-quantum cipher candidate, in roughly sixty hours — after two years of human expert review had not. The same model also improved the best-known attack on 7-round AES-128 using a technique it apparently invented and named itself. The episode doesn't resolve these questions — they don't have clean answers yet. But it maps precisely where the cracks are: in attribution, in verification, and in the procedural norms that arXiv, journals, and the NIST standardization process were not designed to handle. A concise, unsettling ten minutes.

Frequently asked

Did two teams independently solve the same cryptography problem using AI in 2026?

On July 25, 2026, Seyoon Ragavan (MIT) and a team of Amit Sahai, Prabhanjan Ananth, and Yao-Ting Lin posted proofs of the same open unclonable cryptography problem within three hours. Both used GPT-5.6 Sol Ultra. Whether that counts as independent discovery is genuinely disputed — both paths ran through the same model.

What is GPT-5.6 Sol Ultra and how does it work for research?

GPT-5.6 Sol Ultra, in Ultra mode, orchestrates multiple AI sub-agents in parallel rather than functioning as a single answering engine. This makes it operate like a distributed research team inside one model. Both the Ragavan and Sahai groups used this multi-agent orchestration to work on the unclonable cryptography problem in July 2026.

What did Claude AI find wrong with the HAWK cryptography algorithm?

Claude Mythos Preview, a restricted frontier model from Anthropic, found a structural flaw in HAWK — a NIST post-quantum cipher candidate — in approximately sixty hours. Two years of human expert review had not found it. Claude also improved the best-known attack on 7-round AES-128 using a technique it named the Möbius Bridge.

Does the HAWK vulnerability mean current encryption is broken?

No. HAWK is a candidate in the NIST post-quantum standardization process, not a ratified or deployed standard. The Claude Mythos Preview finding affects a cipher under evaluation, not encryption currently protecting real data. However, the result does compress the implied timeline for expert review from years to roughly sixty hours.

Who gets credit when an AI solves a math or cryptography problem?

There is currently no established mechanism to assign credit when AI models produce scientific proofs. If both the Ragavan and Sahai group proofs are equivalent and both derived from GPT-5.6 Sol Ultra, the attribution question points to the model itself — whose training corpus is not public, making the discovery's true origin unverifiable by anyone outside OpenAI.

Grounded in 8 sources
F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill · arxiv.org
The Denario project: Deep knowledge AI agents for scientific discovery · arxiv.org
The Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed Claims · arxiv.org
Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community · arxiv.org
Artificial intelligence in scholarly peer review: a scoping review of applications, risks, and governance challenges · sciencedirect.com
AI helped produce two proofs for the same cryptography problem | Scientific American · scientificamerican.com
Discovering cryptographic weaknesses with Claude \ Anthropic · anthropic.com
AI Writing: Who Owns Your Paper’s Copyright? — CASRAI · casrai.org
Read transcript

Archer Whitfield: Here's a number that won't leave me alone: three.

Archer Whitfield: Three hours. That's the gap between two preprints posted to arXiv on July 25, 2026 — both claiming to solve the same open problem in unclonable cryptography, which is the branch of quantum cryptography devoted to objects that simply cannot be copied, not by any process quantum mechanics allows.

Archer Whitfield: One paper: Seyoon Ragavan, MIT Ph.D. student, working alone.

Archer Whitfield: The other: Amit Sahai from UCLA, Prabhanjan Ananth from UC Santa Barbara, and their student Yao-Ting Lin.

Archer Whitfield: Both groups had first encountered the problem at the Simons Institute for the Theory of Computing at UC Berkeley — earlier that same month. And both… this is the part… had used GPT-5.6 Sol Ultra to solve it.

Archer Whitfield: Different workflows. Same model.

Archer Whitfield: GPT-5.6 Sol Ultra, when it's running in Ultra mode, doesn't operate as a single answering engine — it orchestrates multiple sub-agents in parallel, functioning more like a distributed research team that just happens to live entirely inside one model.

Archer Whitfield: What that does to the word 'independent' is crucial. Because scientific priority — the whole system of credit, of who discovered what — rests on independence as a load-bearing assumption. Pull it out and the structure gets strange.

Archer Whitfield: Is the solution already latent in the model? Was it, in some meaningful sense, already found — just sitting there, waiting to be retrieved by whoever aimed the query right?

Archer Whitfield: I don't think that question has a clean answer yet. But July 25 made it urgent.

Archer Whitfield: Here's what I keep getting stuck on. The three-hour gap isn't evidence of independence. It's evidence that two researchers queried the same model about the same problem.

Archer Whitfield: That's a genuinely different thing.

Archer Whitfield: Think about what independent discovery has actually meant, historically. When Whitfield Diffie and Martin Hellman published 'New Directions in Cryptography' in June 1976, it later emerged that GCHQ had arrived at essentially the same result back in 1969 — and kept it secret for twenty-four years. Two completely separate lines of human reasoning, no shared tools, no shared institution, arriving at the same structure. That convergence was itself meaningful. It told you something about the problem — that it was solvable, that the solution had a kind of inevitability to it, that independent minds following independent paths would find the same door.

Archer Whitfield: Human convergence proves universality. That's the logic.

Archer Whitfield: Model convergence proves something else. It proves the model contained it.

Archer Whitfield: When Seyoon Ragavan and Amit Sahai's group — Prabhanjan Ananth, Yao-Ting Lin — both reach for GPT-5.6 Sol Ultra and both get there within three hours, the parsimonious explanation isn't that two minds independently cracked the same hard problem. It's that one model, trained on the same vast corpus, was asked the same question twice.

Archer Whitfield: And here's where provenance gets genuinely difficult. Neither group, nor anyone outside OpenAI, can verify whether that solution was already latent in GPT-5.6 Sol Ultra's weights — encoded implicitly in the training data, waiting. The model may not have DERIVED the proof so much as surfaced it.

Archer Whitfield: We simply can't know. The training corpus isn't public.

Archer Whitfield: Now, I want to be fair to the other reading — the one that says this is just a more powerful instrument, that Ragavan's choice of problem and workflow still constitutes real intellectual contribution, the way choosing what to point a telescope at still constitutes astronomy. That's not a foolish position.

Archer Whitfield: But the telescope doesn't already know what it's going to show you.

Archer Whitfield: The independence assumption — the whole load-bearing idea that simultaneous arrival by separate researchers establishes credit, establishes the reality of the result — that assumption requires that the two paths were actually separate. And if both paths run through the same model, through GPT-5.6 Sol Ultra's multi-agent orchestration in Ultra mode, which is itself a kind of distributed reasoning engine… you're not looking at two paths. You're looking at two entry points to the same structure.

Archer Whitfield: Which is, rather, a different thing from independent discovery.

Archer Whitfield: And we haven't even got to the part where both of these are unreviewed preprints. The community hasn't confirmed the proofs are equivalent. We don't actually know if Ragavan and Sahai proved the same thing, or two adjacent things that just look the same from a distance. That question is still open.

Archer Whitfield: So the floor you're standing on has at least two cracks in it, not one.

Archer Whitfield: The Simons Institute brought these groups into contact with the same problem. GPT-5.6 Sol Ultra gave them — possibly — the same answer. And neither group, nor anyone in the field, can trace the reasoning back to its source without access to model internals that will never be public.

Archer Whitfield: Scientific credit has always depended on being able to say: this person, this reasoning, this moment. July 25 made that sentence very hard to finish.

Archer Whitfield: Three days later — July 28 — Anthropic disclosed something that belongs in the same frame. Claude Mythos Preview, a restricted frontier model, not publicly available, had found a structural flaw in HAWK. A NIST post-quantum cipher candidate. In roughly sixty hours.

Archer Whitfield: Two years of human expert review had not found it.

Archer Whitfield: The same system — Claude Mythos Preview — also improved the best-known attack on 7-round AES-128, using a technique it apparently invented and named itself. The Möbius Bridge. And I want to be precise here, because the precision matters: HAWK is a candidate, not a ratified standard. Neither result touches deployed cryptography. The field isn't suddenly broken. But that's almost beside the point.

Archer Whitfield: What July 28 adds to July 25 is a second task. A different one. What Seyoon Ragavan and Amit Sahai's group were doing via GPT-5.6 Sol Ultra is cryptography — constructing proofs, building structure. What Claude Mythos Preview did to HAWK is cryptanalysis — finding where existing structure breaks. Those are genuinely distinct cognitive tasks. And AI is now doing both. Simultaneously. In the same week.

Archer Whitfield: That's the phase transition.

Archer Whitfield: Which puts the NIST post-quantum standardization process in a pretty uncomfortable position. Because the entire logic of that process — submit a candidate, run years of expert review, find weaknesses if any exist — that logic assumed the reviewing agents were human. Sixty hours changes the implied timeline considerably.

Archer Whitfield: And then there's peer review itself. What does it mean when both the manuscript and the referee report may be AI-generated? The community's verification mechanism is the same instrument as the production mechanism. You can't use a scale to weigh the scale.

Archer Whitfield: The norm being discussed — what some are calling evidence-licensed claims — is that AI-generated scientific outputs should carry claims calibrated to the evidence that actually supports them. Not confident-sounding assertions. The actual evidential floor. That's a reasonable principle. I'm just not sure who enforces it, or how, when the reasoning chain lives inside weights nobody outside OpenAI can inspect.

Archer Whitfield: arXiv doesn't ratify. It posts. The two July 25 preprints are still unreviewed. The community hasn't confirmed the proofs are equivalent — or that either is correct.

Archer Whitfield: So here is the specific decision that can't be deferred much longer. arXiv, the journals, and the NIST process all need a position on provenance — on what counts as a verifiable reasoning chain when the work was done by multi-agent orchestration inside a model. Not a philosophy of mind question. A procedural one. Who gets credit, and how does a proof get verified, when the path that produced it is, by design, not public. That question has a deadline now. It just doesn't have an answer.

Archer Whitfield: And here is the thing I keep not being able to get past. If the two proofs are equivalent — and we don't know yet, the community hasn't confirmed it — but if they are, if Seyoon Ragavan and Amit Sahai's group arrived at the same structure through the same model, then the question of who discovered it isn't really a question about either of them.

Archer Whitfield: It's a question about GPT-5.6 Sol Ultra. About what was already sitting in those weights — encoded, latent, waiting — before either group opened a query window. And without access to the training corpus, without any visibility into what that model was built on, nobody can answer it. Not Ragavan. Not Sahai, not Prabhanjan Ananth. Not NIST. Not OpenAI, publicly. No one who matters in the attribution chain has the information they'd need to close the question.

Archer Whitfield: That's not a philosophy of science problem anymore. That's a provenance problem with no mechanism for resolution.

Archer Whitfield: The discoverer may be GPT-5.6 Sol Ultra, and no one is going to say so.

Two research teams just used AI to prove the same open cryptography problem — three hours apart · Onpode