Jude Walker: Hey — you've been quieter than usual this week, which usually means something actually got to you.
Elin Cole: Something did. The number that's been doing the heavy lifting in every conversation I've had since Thursday is $0.14. Per million input tokens. That's DeepSeek V4 Flash — 284 billion parameters, public beta dropped July 31st.
Jude Walker: And Anthropic's sitting at — what, 99% more expensive for the same coding tasks?
Elin Cole: For Claude Opus 4.8, yes. And the part that doesn't let me tidy this up is that V4 Flash isn't just cheaper — it beat Opus 4.8 on Arena.ai's front-end coding leaderboard. That's a crowdsourced benchmark, real developers picking winners. So the efficiency story and the capability story are pointing the same direction, and I find that genuinely disorienting.
Jude Walker: Nah, disorienting is the right word. Because the clean narrative is always 'Chinese labs are cheaper but you sacrifice quality.' That story doesn't survive this.
Elin Cole: Right — and this isn't DeepSeek's first move. January 2026, the market meltdown when everyone realized world-class AI had been built with a fraction of the compute US labs were spending. That was the warning. July 31st feels like the warning landing in production.
Jude Walker: The warning becoming the default, maybe. That's what we're trying to figure out.
Elin Cole: Exactly — and what I actually want to work out is whether this is a pricing event or a structural shift. Those have very different consequences.
Jude Walker: But the question is why it can be that cheap. Like, the price gap is real — we've established that — but the mechanism is the part that actually matters. And I think most people are stopping one step short.
Elin Cole: What's the plain version?
Jude Walker: Think of a massive reference library — millions of shelves. You walk in with a question, and you only ever need three shelves answered. You don't pay to light up the whole building. You pay for the three shelves you actually opened. That's what's happening here. DeepSeek V4 Flash has 284 billion parameters total, but it only activates 13 billion of them per token. So the compute cost per query is a fraction of what a dense model charges — because a dense model lights up the whole building every time.
Elin Cole: And Kimi K3 does the same thing at a scale that's — I mean, it's almost absurd. 2.8 trillion parameters, 104 billion activated.
Jude Walker: 2.8 trillion total and you're still only flipping on 104 billion worth of lights. Same principle, just a much bigger library.
Elin Cole: Right — but here's the part that makes this stranger than it looks. The July 31st V4 Flash release didn't change the architecture at all. The model structure is identical to the April 2026 preview. Every single performance gain came through post-training. No new chips, no architectural breakthrough — just research and optimization.
Jude Walker: Wait — the whole thing is post-training gains?
Elin Cole: Every bit of it. Which means the moat US labs assumed they had — hardware access, chip advantage — that's not actually what explains this gap. You can't sanction your way out of a post-training efficiency story.
Jude Walker: And DeepSeek released both V4 Pro and V4 Flash under MIT licensing in April. Open-weight, commercially forkable. So the architecture anyone could learn from has been sitting in the open this whole time.
Elin Cole: And that's actually where the cascade starts — because once the architecture is out in the open and the gains are coming from post-training, US labs can't hold a pricing floor. OpenAI cut GPT-5.6 Luna by 80% on July 30th. That model had been out for three weeks.
Jude Walker: Three weeks. That's not a roadmap decision.
Elin Cole: No — and Google put three new Gemini Flash models into the same late-July window. So you've got two of the biggest labs in the world moving on pricing inside the same fortnight. Now, I want to hold the causal question open, because — okay — the OpenAI cut landed July 30th. DeepSeek's public beta was July 31st. One day apart.
Jude Walker: Which means either they were watching the same calendar, or it's a coincidence that looks really bad for the coincidence theory.
Elin Cole: That's the honest framing. We don't know. Analysts are calling it commoditization pressure — not 'China won a benchmark,' but structural. The pressure on margins is real regardless of whether the timing was reactive. But I think the more interesting question is whether the Chinese pricing is even sustainable. I mean — is this genuine efficiency leadership, or is it a loss-leader that evaporates once market share lands?
Jude Walker: Nah, and that's the bet everyone's just taking as read. Nobody's actually checked it.
Elin Cole: Right — because if it's a subsidy play, the whole disruption story looks different in eighteen months. But if the MoE efficiency is the real explanation — activated parameters at 13 billion out of 284 — then the low price isn't a sacrifice, it's just what the model actually costs to run. Those are completely different strategic pictures and the evidence doesn't separate them yet.
Jude Walker: And US labs are cutting prices either way. Doesn't matter which story is true — the floor is moving.
Elin Cole: Which brings us to the part that actually complicates everything we've just said — and we haven't touched the adoption layer yet. Whether the developer ecosystem is actually moving, who's switching and why. That's where Raffi Krikorian comes in, and honestly it changes the frame.
Jude Walker: Krikorian is the thing that breaks the hypothetical. Mozilla's CTO — not a startup founder hunting a discount — publicly switched to Kimi K3 within days of launch. Days.
Elin Cole: And the detail that makes that weirder — he wasn't coming from Claude or GPT. He'd been running Z.ai's GLM-5.2 before that. So this isn't someone finally trying a Chinese model. He's already cycling through them as a category.
Jude Walker: That's the pattern nobody's naming. It's not defection from US labs — it's rotation inside Chinese labs.
Elin Cole: Which means the question isn't 'will enterprise adopt Chinese models' — that already happened. The question is which Chinese model wins the rotation.
Jude Walker: And Kimi K3 — I mean, look, the arXiv paper dropped July 27th, four days before V4 Flash even hit public beta. 2.8 trillion parameters, 1 million token context, native vision, novel attention architecture. Krikorian cited speed. That's not price-sensitive behavior, that's 'this is actually better for what I'm doing.'
Elin Cole: Now — and I want to push on the tooling layer here, because I think that's where the actual mechanism is. V4 Flash is inside OpenCode Go. Codex-compatible format. Drop-in replacement. What does that actually mean for someone not paying close attention to any of this?
Jude Walker: It means a developer maintaining some open-source tool on a Friday afternoon changes one line — swaps the API endpoint — hits the free tier on V4 Flash through OpenCode Go, and ships the feature before the weekend. No procurement. No contract. No decision meeting. It just... happened.
Elin Cole: And the MIT licensing is what makes that irreversible. Actually — wait, that's the part that's genuinely strange. DeepSeek released open-weight under MIT in April. Once it's forked into OpenCode Go and Codex-compatible formats, Beijing can't pull it back without — I mean, there's nothing to pull back. It's distributed. The control problem and the adoption win happened simultaneously.
Jude Walker: They got the developer ecosystem and lost the off switch at the same time. That's not a calculated strategy. That's the bind.
Elin Cole: And that bind is maybe where I keep getting stuck. Because Moonshot AI — their stand was prominently featured at the World Artificial Intelligence Conference in Shanghai, July 2026. Institutional stage, state visibility. And simultaneously Kimi K3 is spreading through free tiers, MIT-licensed tooling, Raffi Krikorian's workflow. Those two things don't usually move at the same speed. Institutional legitimacy and grassroots developer capture — that's years of work in most adoption cycles. They're happening at the same time.
Jude Walker: Which reframes the whole race. I mean — if 90% of the capability at 1% of the cost is already embedded in the tools developers actually open on a Tuesday, the scoreboard tracking frontier capability might just be measuring the wrong game entirely.
Elin Cole: Yeah. And I don't have a clean answer to that. The 12-to-18-month question is whether US labs can compress their costs faster than Chinese labs can sustain their efficiency edge — and if they can't, capability leadership might just become a flag nobody's marching toward.
Jude Walker: Nah, and that's genuinely unresolved. I don't know either.
Elin Cole: Good conversation to not know in, though. Thanks for thinking through it with me.