Onpode
Cover art for OpenAI, Anthropic, and Google are all rolling out new models in days—here's what they reveal

OpenAI, Anthropic, and Google are all rolling out new models in days—here's what they reveal

October 7, 2026 · 10 min

Eliza Ward & Brian Reed

Google Gemini 4 Argon's introductory price of $2/$10 per million tokens is set to double to $4/$20, and independent analysis found it burns 2.3x more tokens per task — meaning developers who migrate during the promo window may pay more per task than before once the promotional rate expires.

In the final week of September and first days of October 2026, Anthropic, Google, and OpenAI each made significant AI announcements in rapid succession, compressing what would normally be months of competitive news into a single week.

0:0010:03
Get the next episode on Technology →

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Technology →

About this episode

Three AI labs dropped major releases in the same week — Anthropic with Opus 5.5 and Sonnet 5.5, Google with Gemini 4 Argon, OpenAI with the Decisions API — and the coverage mostly chased the sticker prices. This episode is about what the sticker prices don't tell you. The Argon story is the clearest example. The $2/$10 per million token rate looked like a decisive win in the cost race, but it's an introductory price set to double once the promotional window closes. Layered on top of that: independent analysis suggesting Argon burns 2.3x more tokens per task than comparable runs, which erodes the advantage before the promo even ends. The episode works through what that actually means for a developer mid-migration when the math changes underneath them. On benchmarks, Argon wins 12 of 18 categories overall — but Google led with coding, which is specifically where outside analysis found it underperforms. And on AA-Omniscience accuracy, there's a 13-point gap versus GPT-6 Astra that reads, per Latent Space, as a deliberate engineering tradeoff: fewer hallucinations, but a narrower knowledge base. The Anthropic sequencing is quietly the more consequential story. Sonnet 5.5 landed within two percentage points of Opus 5.5 on GDPval-AA — Anthropic's own economically-weighted benchmark — at half the cost. That collapses the case for routing complex work to the flagship for most teams. OpenAI didn't release a model. They released infrastructure for routing among models. Whether that's ahead of the problem or just ahead of the pitch, the documentation hasn't answered yet.

Frequently asked

Is Google Gemini 4 Argon actually cheaper than GPT-6 Astra?

Google Gemini 4 Argon's introductory rate of $2/$10 per million tokens looks 60% cheaper, but the rate is set to double to $4/$20 after the promotional window. Independent analysis by DEV Community writer max_quimby found Argon burns 2.3x more tokens per task, which erases the price advantage even before the promo expires.

How does Claude Sonnet 5.5 compare to Opus 5.5 on benchmarks?

Claude Sonnet 5.5 scores within two percentage points of Opus 5.5 on GDPval-AA, Anthropic's economically-weighted benchmark, at roughly half the cost and 30% faster than Sonnet 5. That narrow gap has led developers to route complex planning tasks to Opus 5.5 while using Sonnet 5.5 for bug fixes and terminal agents.

What is Gemini 4 Argon's accuracy compared to GPT-6 Astra?

Gemini 4 Argon scores 50% on AA-Omniscience accuracy versus 63% for GPT-6 Astra, a 13-point gap. Latent Space frames this as a deliberate engineering tradeoff — Argon hallucinates less but knows less — a choice Google did not prominently disclose in its release materials.

What is OpenAI's Decisions API and how does it work?

OpenAI's Decisions API is infrastructure for routing tasks among AI models based on cost, latency, and capability — not a new model itself. As of launch week, documentation was notably thinner than competing Anthropic and Google releases, leaving the product's real-world routing capability unconfirmed for enterprise buyers.

Which Claude model should I use for AI agents — Opus 5.5 or Sonnet 5.5?

Developer community analysis, including routing logic from users like @neil_xbt, points to Opus 5.5 for complex migrations and planning, and Sonnet 5.5 for bug fixes and terminal agents. With Sonnet 5.5 scoring within two points of Opus 5.5 on GDPval-AA at roughly half the cost, the flagship case has narrowed significantly for most everyday tasks.

Grounded in 9 sources
Anthropic and OpenAI launch cheaper models ↗ · cnbc.com
What slowdown? OpenAI, Anthropic release dueling models as AI price wars heat up | Fortune ↗ · fortune.com
Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models — The AI Daily Brief ↗ · aidailybrief.ai
Which Claude Model for AI Agents? Opus, Sonnet, Haiku ↗ · beam.ai
Frontier AI Models: Live Top 10 Rankings, Evidence and Pricing (October 2026) | BenchLM.ai ↗ · benchlm.ai
Gemini vs Claude vs GPT: Six-Model Comparison (2026) | CallMissed ↗ · callmissed.com
Gemini 4 Argon: Benchmarks, Features, Price and API Access - CometAPI ↗ · cometapi.com
Claude Sonnet 5.5 vs Sonnet 5: What Changed ↗ · cosmicjs.com
Gemini 4 Argon Wins 12 of 18 Benchmarks. Code Isn't One. - DEV Community ↗ · dev.to
Read transcript

Eliza Ward: Hey — good to be here. What a week to be paying attention.

Brian Reed: Right? I've been staring at my notes and I keep — let me see, how do I put this — I keep losing count of how many things actually happened.

Eliza Ward: Three labs in one week. Anthropic dropped Opus 5.5, then Sonnet 5.5 days later — September 28th was the Sonnet date. Google announced Gemini 4 Argon. OpenAI launched the Decisions API. That's — wait, that's a lot.

Brian Reed: The Argon price is the number that stuck with me. Two dollars per million input tokens, ten dollars output. That's the introductory rate Google is leading with.

Eliza Ward: Introductory. That word is doing a lot of work.

Brian Reed: That's exactly what I mean — so think about it like airline pricing. You see a fare Tuesday morning, it looks cheap, you screenshot it. But that fare expires. And then however many bags you're checking — that determines whether the ticket was ever actually cheaper in the first place. The sticker isn't the cost.

Eliza Ward: Right — and the $2/$10 rate is listed to double to $4/$20 after the promotional window closes. Which barely got mentioned in any of the coverage.

Brian Reed: So a developer comparing prices this week is making a decision based on numbers that may not exist in a month. That's — yeah, that's the thing to hold onto.

Eliza Ward: And that's before you even get to the token consumption problem — because the doubling is actually the second hit, not the first.

Brian Reed: What do you mean, the second hit?

Eliza Ward: There's independent analysis — DEV Community, a writer called max_quimby — finding that Argon burns 2.3 times more tokens per task than a comparable run. So even at the $2/$10 introductory rate, you're not actually saving 60%. The per-task math is already eroding before the promotional period ends.

Brian Reed: Hold on — 2.3x. That's not a rounding error. That's the price advantage gone.

Eliza Ward: Well — okay, I want to be careful here. That's third-party, not Google's number. Max_quimby ran specific workloads; we don't know if they generalize. But if even half of that holds, GPT-6 Astra is cheaper per task in some scenarios at regular pricing. That's the independent analysis conclusion, not mine.

Brian Reed: Right, so hedge it — but the direction of the effect is real enough to matter. And here's what makes it concrete: imagine a payment-fraud engineer at a mid-size fintech. Tuesday morning they see Argon at what looks like 60% less than what they're paying. They start a migration. But by the time that migration is actually done — I mean, enterprise migrations aren't overnight — the promo rate has expired. Now they're at $4/$20 and burning 2.3x the tokens. They're paying more per transaction review than they were on Astra.

Eliza Ward: And xAI dropped Grok 4.7 the same week at aggressive pricing, so that engineer now has three numbers to model — not two. The cost comparison just got harder.

Brian Reed: Which — wait, that's actually the thing that doesn't get said. The $2/$10 headline made it look like Google had won the cost conversation. But Grok 4.7 undercuts on price too, even trailing on benchmarks. So the sticker-price race is already three-way and the promo window hasn't even closed yet.

Eliza Ward: The signal here is: don't migrate off a confirmed cost model chasing a promotional number. That's what's actually new — not that Argon is cheap, but that cheap has an expiration date and a hidden multiplier.

Brian Reed: But here's what that framing misses — the cost story got all the oxygen and the benchmark story is actually weirder. Because the take circulating is that Argon won the benchmark race this week. And that's — I mean, it's not wrong exactly, but it's not right either.

Eliza Ward: Twelve of eighteen. That's the number. Argon wins 12 of 18 benchmarks overall — that's real. But Google led with coding. DeepSWE v1.1, 77.9%. That was the headline claim.

Brian Reed: And that's exactly where independent analysis says it underperforms.

Eliza Ward: Right. So Google picked the coding story, and coding is specifically the category where — wait, that's not a small problem. That's the lead claim failing on its own terms.

Brian Reed: The part I don't get is — okay so Latent Space put a number on something that makes this concrete. Argon scores 50% on AA-Omniscience accuracy. Astra is at 63%.

Eliza Ward: Fifty versus sixty-three. That's — hold on, that's not a rounding difference. That's thirteen points.

Brian Reed: And Latent Space frames it as a deliberate engineering tradeoff — Argon hallucinates less but knows less. Google chose that. They just didn't, you know, put it in the press release.

Eliza Ward: Which connects to the OSS bug-bounty pause — Google paused that program because AI-generated noise was overwhelming it. And skeptics immediately pointed at that as undercutting the cybersecurity defense claims for Argon. I'm not saying it's a smoking gun, but you can't lead with cybersecurity capability and simultaneously be pausing the program that would surface bugs in it. That's a credibility question worth naming.

Brian Reed: No single model beats Astra across every dimension — that's actually the finding. And the routing question that falls out of this, the workload-by-task logic, is where the Anthropic sequencing gets really interesting — we'll get to how Sonnet 5.5 landing within two points of Opus 5.5 on GDPval-AA basically collapses the flagship argument for most workloads.

Eliza Ward: Two points. On GDPval-AA — Anthropic's own economically-weighted benchmark — Sonnet 5.5 lands within two percentage points of Opus 5.5. That's the gap. That's what the sequencing actually exposed.

Brian Reed: Which means — hang on, let me think through what that actually does to a buyer. If I'm a team lead, I've been routing my most complex planning work to the flagship because that's what you pay flagship prices for. And now Anthropic is telling me the everyday model is two points behind on the benchmark that's supposed to measure real knowledge-work value.

Eliza Ward: At roughly half the cost.

Brian Reed: Right — and 30% faster than Sonnet 5 at the same price. So the developer community discourse on X, the routing logic people like @neil_xbt were laying out — Opus 5.5 for migrations and planning, Sonnet 5.5 for bug fixes and terminal agents — that's not preference, that's actual math now. The flagship case collapses for most everyday tasks.

Eliza Ward: Wait — does it collapse or does it just narrow? Because Dario Amodei's pacing pledge is sitting in the background of all this, and I want to name it without overstating it. Anthropic ran a five-model streak. That's confirmed. The exact terms of the pledge? Not clearly documented. So I can't press that irony as hard as I want to.

Brian Reed: The part I don't get is — if you're Anthropic and you've just collapsed the performance gap between your flagship and your mid-tier model yourself, are you confident that's a ceiling, or are you fragmenting defensively because Google and OpenAI are both in the room this week?

Eliza Ward: Genuinely don't know. And that's where the Decisions API lands differently — because OpenAI's move isn't a new model at all. It's infrastructure for routing among models by task, latency, cost, capability. If that works, OpenAI benefits from the fragmentation instead of losing to it.

Brian Reed: That's — actually, that reframes the whole week. But the documentation on the Decisions API was notably thinner than either the Anthropic or Google releases. So I genuinely cannot tell if OpenAI is ahead of this problem or just announcing before they've figured out the pitch.

Eliza Ward: That's the honest answer. Watch for actual Decisions API documentation — that's the concrete signal. Until that lands, the routing infrastructure story is real as a concept and thin as a product.

Brian Reed: The window closing on Argon's pricing — that's the concrete moment. When $2/$10 becomes $4/$20, we'll actually know whether the real-world cost argument holds or just... dissolves.

Eliza Ward: That's the date to watch. And if it dissolves — I mean, the 2.3x token burn is already in the background, so at $4/$20 with that multiplier, the math that made Argon interesting this week basically doesn't exist anymore. Then what's left? Twelve of eighteen benchmarks and a 13-point accuracy gap on AA-Omniscience.

Brian Reed: Which is where the Decisions API question lands — wait, actually this is the part I genuinely can't answer. If that routing infrastructure gets real documentation and starts working, OpenAI absorbs the fragmentation. If it stays thin, no one has a clean answer for the enterprise buyer once Argon's promo rate expires.

Eliza Ward: And whether Sonnet 5.5 genuinely replaces Opus 5.5 for most workloads — that's not settled either. The GDPval-AA gap is two points, but that's Anthropic's number on Anthropic's benchmark.

Brian Reed: None of it resolves this week.

OpenAI, Anthropic, and Google are all rolling out new models in days—here's what they reveal · Onpode