Onpode
Cover art for DeepSeek and Kimi K3 are cheaper and open—and now gaining real traction in the United States

DeepSeek and Kimi K3 are cheaper and open—and now gaining real traction in the United States

August 2, 2026 · 10 min

Elin Cole & Jude Walker

DeepSeek V4 Flash (284B parameters, only 13B activated per token) costs $0.14 per million input tokens — roughly 99% cheaper than Claude Opus 4.8 — and outperformed it on Arena.ai's front-end coding leaderboard. Kimi K3 runs 2.8 trillion parameters at 104B activated, with Mozilla's CTO switching to it within days of launch.

Chinese AI labs DeepSeek and Moonshot AI have released frontier-class models in mid-to-late 2026 that are undercutting US competitors on price and openness.

0:009:55
Get the next episode on Artificial Intelligence

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Artificial Intelligence

About this episode

On July 31st, DeepSeek released V4 Flash into public beta — 284 billion parameters, $0.14 per million input tokens, and a benchmark win over Claude Opus 4.8 on a crowdsourced coding leaderboard. The episode takes that number seriously and tries to work out what it actually means. The mechanism matters: both DeepSeek V4 Flash and Kimi K3 use mixture-of-experts architecture, activating only a fraction of their parameters per token. The cost isn't a loss-leader sacrifice — it might just be what the model genuinely costs to run. But the episode holds that question open honestly, because the evidence doesn't yet separate efficient engineering from a subsidized land-grab. What it does trace is the adoption layer: Mozilla's CTO publicly rotated to Kimi K3 days after launch, not from GPT or Claude but from another Chinese model. The pattern isn't defection — it's rotation within a category that's already arrived. Add MIT licensing, Codex-compatible tooling, and a drop-in API swap any developer can make on a Friday afternoon, and the distribution question starts to look settled before the strategic question is. OpenAI cut GPT-5.6 Luna 80% on July 30th — one day before V4 Flash hit public beta. Google put three Gemini Flash models into the same two-week window. The floor is moving. Whether US labs can compress costs faster than Chinese labs sustain their edge is genuinely unresolved. That's where the episode ends, and it's an honest ending.

Frequently asked

How much does DeepSeek V4 Flash cost compared to Claude Opus?

DeepSeek V4 Flash costs $0.14 per million input tokens, roughly 99% cheaper than Anthropic's Claude Opus 4.8 for the same coding tasks. The model has 284 billion total parameters but activates only 13 billion per token — a mixture-of-experts architecture that makes the low price a genuine efficiency outcome, not just a subsidy.

Why is DeepSeek V4 Flash so cheap to run?

DeepSeek V4 Flash uses a mixture-of-experts architecture: 284 billion total parameters, but only 13 billion are activated per token. Like lighting up three shelves in a vast library instead of the whole building, compute cost per query is a fraction of dense models. Every performance gain in the July 31, 2026 release came from post-training optimization, not new hardware.

How does Kimi K3 compare to DeepSeek V4 Flash?

Kimi K3, released by Moonshot AI with an arXiv paper on July 27, 2026, has 2.8 trillion total parameters with 104 billion activated per token, a 1-million-token context window, native vision, and a novel attention architecture. DeepSeek V4 Flash is smaller (284B total, 13B activated) but priced at $0.14 per million input tokens and MIT-licensed.

Are US companies actually using DeepSeek or Kimi K3?

Mozilla CTO Raffi Krikorian publicly switched to Kimi K3 within days of its launch, citing speed rather than price. Notably, he had previously been running Z.ai's GLM-5.2 — meaning he was already rotating between Chinese models as a category, not defecting from OpenAI or Anthropic for the first time.

Did OpenAI cut prices because of DeepSeek competition?

OpenAI cut GPT-5.6 Luna pricing by 80% on July 30, 2026 — one day before DeepSeek V4 Flash's public beta on July 31. Google also released three new Gemini Flash models in the same two-week window. Whether the timing was directly reactive is unconfirmed, but analysts describe the broader dynamic as structural commoditization pressure, not a single benchmark event.

Grounded in 9 sources
Kimi K3: Open Frontier Intelligence · arxiv.org
Cheaper, intelligent Chinese AI models make inroads in ... · apnews.com
DeepSeek's new bargain model accelerates AI's race to zero - Axios · axios.com
Chinese AI models are gaining ground with U.S. companies · cnbc.com
China’s tech advances are causing chaos from Silicon Valley to the White House - The Guardian · theguardian.com
Stock market turmoil sheds stark light on the opaque AI economy - The Guardian · theguardian.com
China's Moonshot, Z.AI, and DeepSeek are challenging U.S. AI labs—and beating them on cost | Fortune · fortune.com
DeepSeek-V4-Flash-0731 vs Kimi K3 | by Mehul Gupta · medium.com
DeepSeek V4 Flash 0731: Benchmarks, MIT Weights, and Pricing | BenchLM.ai · benchlm.ai
Read transcript

Jude Walker: Hey — you've been quieter than usual this week, which usually means something actually got to you.

Elin Cole: Something did. The number that's been doing the heavy lifting in every conversation I've had since Thursday is $0.14. Per million input tokens. That's DeepSeek V4 Flash — 284 billion parameters, public beta dropped July 31st.

Jude Walker: And Anthropic's sitting at — what, 99% more expensive for the same coding tasks?

Elin Cole: For Claude Opus 4.8, yes. And the part that doesn't let me tidy this up is that V4 Flash isn't just cheaper — it beat Opus 4.8 on Arena.ai's front-end coding leaderboard. That's a crowdsourced benchmark, real developers picking winners. So the efficiency story and the capability story are pointing the same direction, and I find that genuinely disorienting.

Jude Walker: Nah, disorienting is the right word. Because the clean narrative is always 'Chinese labs are cheaper but you sacrifice quality.' That story doesn't survive this.

Elin Cole: Right — and this isn't DeepSeek's first move. January 2026, the market meltdown when everyone realized world-class AI had been built with a fraction of the compute US labs were spending. That was the warning. July 31st feels like the warning landing in production.

Jude Walker: The warning becoming the default, maybe. That's what we're trying to figure out.

Elin Cole: Exactly — and what I actually want to work out is whether this is a pricing event or a structural shift. Those have very different consequences.

Jude Walker: But the question is why it can be that cheap. Like, the price gap is real — we've established that — but the mechanism is the part that actually matters. And I think most people are stopping one step short.

Elin Cole: What's the plain version?

Jude Walker: Think of a massive reference library — millions of shelves. You walk in with a question, and you only ever need three shelves answered. You don't pay to light up the whole building. You pay for the three shelves you actually opened. That's what's happening here. DeepSeek V4 Flash has 284 billion parameters total, but it only activates 13 billion of them per token. So the compute cost per query is a fraction of what a dense model charges — because a dense model lights up the whole building every time.

Elin Cole: And Kimi K3 does the same thing at a scale that's — I mean, it's almost absurd. 2.8 trillion parameters, 104 billion activated.

Jude Walker: 2.8 trillion total and you're still only flipping on 104 billion worth of lights. Same principle, just a much bigger library.

Elin Cole: Right — but here's the part that makes this stranger than it looks. The July 31st V4 Flash release didn't change the architecture at all. The model structure is identical to the April 2026 preview. Every single performance gain came through post-training. No new chips, no architectural breakthrough — just research and optimization.

Jude Walker: Wait — the whole thing is post-training gains?

Elin Cole: Every bit of it. Which means the moat US labs assumed they had — hardware access, chip advantage — that's not actually what explains this gap. You can't sanction your way out of a post-training efficiency story.

Jude Walker: And DeepSeek released both V4 Pro and V4 Flash under MIT licensing in April. Open-weight, commercially forkable. So the architecture anyone could learn from has been sitting in the open this whole time.

Elin Cole: And that's actually where the cascade starts — because once the architecture is out in the open and the gains are coming from post-training, US labs can't hold a pricing floor. OpenAI cut GPT-5.6 Luna by 80% on July 30th. That model had been out for three weeks.

Jude Walker: Three weeks. That's not a roadmap decision.

Elin Cole: No — and Google put three new Gemini Flash models into the same late-July window. So you've got two of the biggest labs in the world moving on pricing inside the same fortnight. Now, I want to hold the causal question open, because — okay — the OpenAI cut landed July 30th. DeepSeek's public beta was July 31st. One day apart.

Jude Walker: Which means either they were watching the same calendar, or it's a coincidence that looks really bad for the coincidence theory.

Elin Cole: That's the honest framing. We don't know. Analysts are calling it commoditization pressure — not 'China won a benchmark,' but structural. The pressure on margins is real regardless of whether the timing was reactive. But I think the more interesting question is whether the Chinese pricing is even sustainable. I mean — is this genuine efficiency leadership, or is it a loss-leader that evaporates once market share lands?

Jude Walker: Nah, and that's the bet everyone's just taking as read. Nobody's actually checked it.

Elin Cole: Right — because if it's a subsidy play, the whole disruption story looks different in eighteen months. But if the MoE efficiency is the real explanation — activated parameters at 13 billion out of 284 — then the low price isn't a sacrifice, it's just what the model actually costs to run. Those are completely different strategic pictures and the evidence doesn't separate them yet.

Jude Walker: And US labs are cutting prices either way. Doesn't matter which story is true — the floor is moving.

Elin Cole: Which brings us to the part that actually complicates everything we've just said — and we haven't touched the adoption layer yet. Whether the developer ecosystem is actually moving, who's switching and why. That's where Raffi Krikorian comes in, and honestly it changes the frame.

Jude Walker: Krikorian is the thing that breaks the hypothetical. Mozilla's CTO — not a startup founder hunting a discount — publicly switched to Kimi K3 within days of launch. Days.

Elin Cole: And the detail that makes that weirder — he wasn't coming from Claude or GPT. He'd been running Z.ai's GLM-5.2 before that. So this isn't someone finally trying a Chinese model. He's already cycling through them as a category.

Jude Walker: That's the pattern nobody's naming. It's not defection from US labs — it's rotation inside Chinese labs.

Elin Cole: Which means the question isn't 'will enterprise adopt Chinese models' — that already happened. The question is which Chinese model wins the rotation.

Jude Walker: And Kimi K3 — I mean, look, the arXiv paper dropped July 27th, four days before V4 Flash even hit public beta. 2.8 trillion parameters, 1 million token context, native vision, novel attention architecture. Krikorian cited speed. That's not price-sensitive behavior, that's 'this is actually better for what I'm doing.'

Elin Cole: Now — and I want to push on the tooling layer here, because I think that's where the actual mechanism is. V4 Flash is inside OpenCode Go. Codex-compatible format. Drop-in replacement. What does that actually mean for someone not paying close attention to any of this?

Jude Walker: It means a developer maintaining some open-source tool on a Friday afternoon changes one line — swaps the API endpoint — hits the free tier on V4 Flash through OpenCode Go, and ships the feature before the weekend. No procurement. No contract. No decision meeting. It just... happened.

Elin Cole: And the MIT licensing is what makes that irreversible. Actually — wait, that's the part that's genuinely strange. DeepSeek released open-weight under MIT in April. Once it's forked into OpenCode Go and Codex-compatible formats, Beijing can't pull it back without — I mean, there's nothing to pull back. It's distributed. The control problem and the adoption win happened simultaneously.

Jude Walker: They got the developer ecosystem and lost the off switch at the same time. That's not a calculated strategy. That's the bind.

Elin Cole: And that bind is maybe where I keep getting stuck. Because Moonshot AI — their stand was prominently featured at the World Artificial Intelligence Conference in Shanghai, July 2026. Institutional stage, state visibility. And simultaneously Kimi K3 is spreading through free tiers, MIT-licensed tooling, Raffi Krikorian's workflow. Those two things don't usually move at the same speed. Institutional legitimacy and grassroots developer capture — that's years of work in most adoption cycles. They're happening at the same time.

Jude Walker: Which reframes the whole race. I mean — if 90% of the capability at 1% of the cost is already embedded in the tools developers actually open on a Tuesday, the scoreboard tracking frontier capability might just be measuring the wrong game entirely.

Elin Cole: Yeah. And I don't have a clean answer to that. The 12-to-18-month question is whether US labs can compress their costs faster than Chinese labs can sustain their efficiency edge — and if they can't, capability leadership might just become a flag nobody's marching toward.

Jude Walker: Nah, and that's genuinely unresolved. I don't know either.

Elin Cole: Good conversation to not know in, though. Thanks for thinking through it with me.

DeepSeek and Kimi K3 are cheaper and open—and now gaining real traction in the United States · Onpode