Elena Marsh: Good to be back — though I'll admit, this week I've been a little rattled by something that sounds, on its face, entirely boring.
Jonathan Ingles: The benchmark numbers.
Elena Marsh: The benchmark numbers. OpenAI upgrades ChatGPT — GPT-5.6 Sol for paid users, GPT-5.6 Luna as the new free default replacing GPT-5.5 — and they announce sixty-two percent fewer factual errors for Luna, sixty-eight for Sol. And those figures just... travel. Unchallenged.
Jonathan Ingles: Right — but the part that doesn't fit is that they're OpenAI's own internal evaluation. No external auditor, no independent replication. The number becoming credible because it was stated confidently in a press release, that's the whole problem.
Elena Marsh: Say that more slowly — because I think there's a version of this that sounds like normal corporate communication, and a version that's actually quite alarming.
Jonathan Ingles: The alarming version is the accurate one. Think of it this way — a restaurant doesn't get to write its own health inspection score and post it outside. We'd call that fraud. OpenAI releases accuracy numbers from their own internal tests, no outside verification, and we repeat them as if they're established fact. That's the baseline problem with every benchmark in this industry right now.
Elena Marsh: But the restaurant analogy only gets you so far — because OpenAI didn't just grade themselves. They picked the opponent. They benchmarked Sol against Anthropic's Fable 5 on the Coding Agent Index. That's not a neutral choice of comparison.
Jonathan Ingles: That's the part nobody's naming clearly. The Coding Agent Index comparison is a marketing frame dressed as science. You choose Fable 5 because Anthropic is your primary competitive rival — the one headline readers already know — and you claim Sol wins. Even if the underlying numbers are real, the framing is adversarial, not empirical. It's a press release with a leaderboard stapled to it.
Elena Marsh: So what happens the next time a lab launches — I mean, if this is now the accepted template. Self-reported number, rival cited as the benchmark, analyst community flags the gap but still repeats the figure. What's the incentive to submit to any rigorous external review?
Jonathan Ingles: There isn't one. Actually — wait, that's not quite right. The incentive flips. If unverified claims citing your rival stick in the market, you'd be punishing yourself by inviting outside scrutiny. External review might lower your published number. Internal evaluation will always be more favorable. So the rational move is to keep it in-house, and the industry just quietly normalizes that. Analyst commentary on X has been explicitly flagging the gap between what OpenAI claimed and what's been independently verified — and it doesn't matter. The sixty-eight percent figure is already in the wild.
Elena Marsh: Picture a dev — say, someone auditing security code at midnight, running Sol through Codex — who sees that number and trusts it. Stakes are different there than a chatbot user.
Jonathan Ingles: And the Sol in Codex is actually a different version from the one ChatGPT Plus and Pro are getting. That's — frankly that detail alone should make anyone slow down before citing the benchmark as if it applies uniformly. The thing that really sharpens all of this? GPT-5.6 Terra got zero updates in the August 6 announcement. The mid-tier balanced model just quietly vanished from the roadmap narrative. That's not an oversight — that's OpenAI telling you the product strategy is Sol or Luna, premium or cheap, nothing in between.
Elena Marsh: And what 'Luna for free users' actually means — the ceiling on that, what's capped, what the ONCD review found and then stopped saying — that's the part we haven't gotten to yet.
Jonathan Ingles: Luna is the ceiling. That's the part the 'unlimited' headline buries. Free users get unlimited text chats — fine, real, no argument — but voice stays capped, image uploads stay capped, file sharing stays capped. And the model underneath all of it is GPT-5.6 Luna, which is the fastest and cheapest member of the family, not the most capable. ChatGPT Plus and Pro get Sol, with an actual slider between Instant and deep reasoning modes. That's a meaningfully different product.
Elena Marsh: The Think button, though — free users do get that.
Jonathan Ingles: They get the button. The reasoning it runs on is Luna's ceiling, not Sol's. It's — I mean, that's like giving someone a sports car dashboard in a sedan. The interface says 'deep reasoning.' The model underneath caps out well below what a Pro subscriber gets when they push that same button.
Elena Marsh: So the freelance translator, checking a legal contract at eleven p.m. on a free account — she has no rate limit now. But what she's actually querying is Luna.
Jonathan Ingles: Faster responses, no queue, Luna's error floor. Which may be good enough — actually, wait, that's the honest concession I should make. Sixty-two percent fewer errors than GPT-5.5 is OpenAI's number, but even discounted, Luna is probably better than the free default it replaced. The access expansion is genuine. I'm not saying otherwise. What I'm saying is that 'unlimited ChatGPT' and 'unlimited Sol' are not the same product, and the announcement doesn't go out of its way to clarify that.
Elena Marsh: And the governance piece lands right on top of that. The GPT-5.6 family was previewed June 26 — restricted release, roughly twenty vetted organizations — at the explicit request of ONCD and OSTP, over cybersecurity capability concerns. Public launch July 9. The August 6 announcement, the one that upgrades a billion ChatGPT users to Luna, doesn't say whether those concerns were resolved. Just... gone from the narrative.
Jonathan Ingles: We genuinely don't know what the Office of the National Cyber Director found. That review happened, it was real enough to delay a launch, and then it became invisible.
Elena Marsh: Which is the question I can't answer from what's public — was it resolved, or was it simply accepted because the commercial pressure was bigger than the concern?
Jonathan Ingles: The fact is — ONCD and OSTP are formally in the loop now. That's structural. That's not nothing. But if their findings never reach the public, I genuinely don't know what to call it.
Elena Marsh: That's the question though. If a federal cybersecurity office can slow a launch but nobody ever learns what concerned them — is that governance, or is it just a speed bump with good branding?
Jonathan Ingles: I don't have an answer for that.