Brian Reed: Hi. Okay, I have to ask — did the last two weeks feel normal to you?
Eliza Ward: Not even slightly. GPT-6 Astra drops September 3, and then eleven days later Brockman is on the a16z podcast saying "we are now in the AGI era." Out loud. On tape.
Brian Reed: Hundreds of thousands of views on that clip. Though — and this is the part I keep wanting to flag — he actually said it at the Astra launch briefings too, on September 3. So the podcast wasn't first.
Eliza Ward: No, it wasn't. The September 3 statement existed. The a16z version is when it actually — I mean, it escaped the briefing room.
Brian Reed: And Jensen Huang amplified it on X with his own "AGI has arrived" post, same stretch of days. So at what point is this a coordinated moment versus two people who happen to agree?
Eliza Ward: We don't know that yet. What we know is it's moving.
Brian Reed: But here's where the moving part gets complicated — the number they're moving on might not be real.
Eliza Ward: The ARC-AGI-3 score.
Brian Reed: Ninety-nine point nine percent. That's what OpenAI reported, on their own internal harness. And then an independent team runs the same benchmark, standard conditions, and gets... sixty-two point seven. That's a 37-point gap. I mean — imagine a student who scores a hundred on a practice test their own teacher wrote, then sixty-three when an outside examiner gives the same subject. You wouldn't lead with the hundred.
Eliza Ward: And OpenAI has said nothing about it. No methodology note. No "here's why the conditions differ." Nothing public.
Brian Reed: The silence is — yeah, that's the part that I don't know what to do with.
Eliza Ward: Okay, and then layer on the Humanity's Last Exam number — Astra scored fifty-four point eight percent on that. Which is not — wait, that's not a superhuman score on a hard exam. That's a little above average for the exam, depending on the field. So the internal harness gives you a near-perfect score, the independent one gives you sixty-three, and a separate hard benchmark gives you fifty-five. Those three numbers don't tell the same story.
Brian Reed: And OpenAI's own charter defines AGI as systems that outperform humans at most economically valuable work. Benchmark scores — even the real ones — can't actually prove that threshold.
Eliza Ward: That's the actual gap. Not 37 points on one test — it's that no benchmark score, verified or not, meets the standard OpenAI wrote for itself.
Brian Reed: So what do we need to see? Like, what would close that gap?
Eliza Ward: Closing that gap is actually — wait, that's where the whole thing gets slippery. Because Brockman's answer to 'what would count' is already in the statement. He said 'everyone has a different definition' of AGI. In the same breath as 'we are in the AGI era.' That's not an admission of complexity. That's a trapdoor. You can't falsify a claim that comes with its own escape hatch built in.
Brian Reed: And that's the circulating take I want to push back on — the one that says Pachocki's caution post contradicts the declaration. Like, OpenAI's own chief scientist calling for extreme caution and voluntary slowdowns, same week, same company — people read that as the left hand not knowing what the right hand is doing.
Eliza Ward: It's not a contradiction.
Brian Reed: No, it's — I mean, it's almost the opposite of a contradiction. Pachocki's caution actually reinforces the declaration. If you're calling for extreme caution, you're implicitly confirming something serious has happened. The danger talk and the arrival talk aren't in tension. They're two parts of the same message, just... running at different speeds for different audiences.
Eliza Ward: And Sam Altman ruling out a 2026 IPO citing safety — same week — fits the same pattern. Actually, hold on, that one genuinely surprised me. If AGI has arrived and it's dangerous enough to block a public offering, why would you not raise capital publicly to build defensive infrastructure? The two positions only cohere if the safety framing is also doing narrative work.
Brian Reed: And then the 2026 International AI Safety Report lands in the middle of all this saying current systems show 'early signs of some relevant capabilities' — not loss-of-control levels. That's not a thumbs-up for the AGI-era framing. Amodei, Hassabis, LeCun, Karpathy — they're all pointing at real gaps: continual learning, world models, long-horizon reliability. So the external scientific read and the declaration aren't... they're not describing the same thing.
Eliza Ward: Right — but none of that external skepticism can falsify the claim, because the claim never specified what would falsify it. That's the design.
Brian Reed: Which means the part that actually matters might be who moves on the declaration and on what signal — because someone already has. It gets worse before it gets better.
Eliza Ward: Yeah, Bernie Sanders and Greg Casar dropped the Ban Artificial Superintelligence Act on September 3 — the same day as the Astra launch. Policy is already reacting to a framing that the scientific community hasn't verified. That's the next thread.
Brian Reed: About that — the Act targets superintelligence, not AGI. Those are different words. Sanders and Casar wrote a bill in response to a framing that doesn't even match the term Brockman used.
Eliza Ward: Twenty years in prison for building superintelligent AI. That's the criminal penalty in the Act. And the declaration that moved them was about AGI.
Brian Reed: That's — the term mismatch alone is a problem. Policy is already running ahead of the definition.
Eliza Ward: And here's the actual mechanism. The morning of September 15, a staffer in a Senate office sees the a16z clip headline — 'OpenAI President: We Are in the AGI Era' — and they don't read the transcript. They don't wait for independent verification. They start drafting briefing language around 'AGI arrival.' That's done before any external review has been published. The declaration is already inside the policy process.
Brian Reed: And the briefing language then shapes the next Sanders-Casar conversation, which already used the wrong term. So the error compounds.
Eliza Ward: Which is exactly why Brockman's framing — 'when people look back' — is actually load-bearing here. He didn't say 'this is AGI, verifiable now.' He positioned it as something history will confirm. You can't falsify a retroactive claim in the present tense. So while researchers are debating the 62.7 versus 99.9 gap, the narrative is already—
Brian Reed: It's already in the Senate.
Eliza Ward: The concrete signal I'd watch is Anthropic and Google DeepMind. If either of them ships something comparably capable in the next few months and just — doesn't call it the AGI era — that non-adoption is the clearest evidence the September 14 declaration was positioning. Dario Amodei has already been cautioning against this framing. If his next model is in the same tier and he stays silent on the AGI label, that silence does more to answer the question than any benchmark.
Brian Reed: So the verification test for the declaration isn't a benchmark score. It's whether the rest of the industry ratifies the language.
Eliza Ward: And that's genuinely the cleanest test available. Not another benchmark. Not OpenAI's internal harness. Whether Anthropic or Google DeepMind ship something comparable and just — don't reach for the label. That's it. That's what settles it.
Brian Reed: And we actually have a baseline for that. OpenAI's own charter — the definition they wrote — says AGI means outperforming humans at most economically valuable work. Not one benchmark. Not 99.9% on a harness they control. Most economically valuable work. That standard is just... sitting there, unmet, and nobody at the company has pointed at it and said 'here's where Astra crosses it.' So I mean — I don't know if the September 14 declaration was a technical milestone or positioning. I genuinely don't. But I know which question would answer it, and it's not one Brockman has touched.
Eliza Ward: Wait — do you think the silence from Dario Amodei on the label, if it comes, actually changes anything? Or does the declaration just... hold, regardless?
Brian Reed: That's the part I can't resolve. Actually — no, let me be more specific. I think Amodei's silence would matter to researchers and maybe to some regulators. I don't think it touches the venture capital repricing or the Senate briefing language. Those are already downstream of the declaration. The belief is already doing work whether or not the fact catches up.
Eliza Ward: Which is the open question I don't have an answer to. Can a declaration become true through the act of being believed. We're not there yet. We might be by spring.