Onpode
Cover art for Meta launched Muse Code, a new coding agent priced below Claude and OpenAI's Codex, directly competing in the developer automation market

Meta launched Muse Code, a new coding agent priced below Claude and OpenAI's Codex, directly competing in the developer automation market

August 6, 2026 · 9 min

Eliza Ward & Brian Reed

Meta launched Muse Code on August 5, 2025, priced at $1.25 per million input tokens — undercutting Claude Code and OpenAI's Codex from day one. A 'contributor tier' at $0.10/million is cheaper still, but Meta has not disclosed what codebase data that tier collects, creating real enterprise risk despite the competitive benchmark score of 82.9% on Terminal-Bench 2.1.

On August 5, 2026, Meta Platforms announced Muse Code, a terminal-based AI coding agent now available in public beta, entering direct competition with Anthropic's Claude Code and OpenAI's Codex.

0:009:09
Get the next episode on Claude

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Claude

About this episode

On August 5th, Meta launched Muse Code — a coding agent that drops into your terminal, reads your entire codebase, and handles complex multi-step tasks across large repositories. The headline number is the price: $1.25 per million input tokens, undercutting both Claude Code and OpenAI's Codex on day one. But the more interesting story is the contributor tier, which goes even further — $0.10 per million input tokens, roughly ten times cheaper than standard. The exchange: you provide feedback. What that feedback includes is, as of launch, unspecified. The episode works through what that ambiguity actually means for a developer who runs Muse Code on a payments module Saturday afternoon — before Meta has published any data governance terms at all. It also looks hard at the benchmark claims. Muse Code scores 82.9% on Terminal-Bench 2.1, behind Claude Code's 86.7%, but ahead of Codex and Grok Build. The problem: no source explains what Terminal-Bench 2.1 actually tests, who runs it, or whether the tasks it measures resemble the complex, multi-repo refactors Muse Code is explicitly designed for. And underneath all of it sits a structural shift worth watching — a company that spent years arguing open-source AI is safer has just launched a fully proprietary, cloud-only product. The question the episode ends on: will enterprise teams adopt the contributor tier because the price won, or avoid it because the terms weren't there yet? Those are two different answers to the same question.

Frequently asked

How does Meta Muse Code pricing compare to Claude Code and OpenAI Codex?

Meta Muse Code's standard tier costs $1.25 per million input tokens, versus Claude Code bundled in Anthropic's $20/month Pro plan and Codex through ChatGPT Plus at the same flat rate. A contributor tier drops Muse Code to $0.10 per million input tokens in exchange for user feedback to Meta.

How does Muse Code perform on benchmarks compared to Claude Code?

Muse Spark 1.2, the model behind Muse Code, scores 82.9% on Terminal-Bench 2.1. Claude Code on Opus 5 scores 86.7% — a 3.8-point gap. Codex on GPT-5.6 Terra scores 81.8%. No independent source has audited what Terminal-Bench 2.1 actually tests, so the gap's real-world meaning is unclear.

What is Meta Muse Code's contributor tier and what does it cost?

Meta Muse Code's contributor tier costs $0.10 per million input tokens and $0.20 per million output tokens — roughly ten times cheaper than the standard tier. The price reduction is exchanged for user feedback that Meta has not fully defined, and the company has not disclosed whether that feedback includes codebase contents.

Is Meta Muse Code safe for enterprise or startup codebases?

Meta has not disclosed what data Muse Code's contributor tier collects, how long it is retained, or whether repository contents are included in 'product feedback.' Muse Code is cloud-only — no on-premise option — meaning code does not stay on the user's machine. Enterprise data governance terms had not been published at launch on August 5, 2025.

How does Meta Muse Code differ from Meta's open-source Llama models?

Meta Muse Code is fully proprietary and cloud-only, a structural reversal from Llama, which Meta positioned as open and self-hostable with no data handoff. Muse Code is accessible via Meta's developer site and OpenRouter, but every access path routes through Meta's cloud infrastructure.

Grounded in 8 sources
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems · arxiv.org
Harness Engineering for Agentic AI Coding Tools: An Exploratory Study · arxiv.org
Meta Debuts AI Coding Agent Muse: Here's How It Compares to Claude ... · tech.yahoo.com
Meta to take on Anthropic's Claude and OpenAI's Codex with new coding agent · businessinsider.com
GitHub - NanoNets/Graft: Turbocharge Claude Code, Cursor, Codex, Gemini & every coding agent: faster, cheaper, with contextual understanding specific to your codebase. · GitHub · github.com
Meta enters the AI coding wars with Muse Spark 1.2 and Muse Code with persistent async background agents | VentureBeat · venturebeat.com
https://www.cryptopolitan.com/metas-coding-ai-spark-another-ai-price-war/ · cryptopolitan.com
Meet Muse Spark 1.2 and Muse Code · developer.meta.com
Read transcript

Brian Reed: Hey. You caught the Zuckerberg post this morning?

Eliza Ward: Yeah — the @finkd one. Meta just dropped a coding agent into public beta. Today. August fifth.

Brian Reed: Muse Code. And I mean — I've been sitting with this for like twenty minutes and the thing that won't leave me alone is the pricing. Because Alexandr Wang goes on X and basically says, look, we're not claiming we're the best. We're saying we're the cheapest. That's the pitch.

Eliza Ward: Standard rate is $1.25 per million input tokens. Claude Code — Anthropic's — that's bundled in a $20/month Pro plan. Codex, same thing through ChatGPT Plus.

Brian Reed: So Meta's undercutting both of them on day one.

Eliza Ward: On day one. And Wang is framing that as the feature, not a concession.

Brian Reed: Which — okay, hang on — does that hold? Because Meta Superintelligence Labs also published the Terminal-Bench 2.1 numbers. Muse Spark 1.2 scores 82.9%. Claude Code on Opus 5 is at 86.7. That's nearly four points back. So they're leading with price because... they have to?

Eliza Ward: That's the thing I want to — wait, actually that framing might be too clean. We don't know yet whether four points on Terminal-Bench 2.1 translates to anything a developer would actually feel.

Brian Reed: Right — but the four points thing, I want to set that aside for a second because I think it's actually the wrong frame. The real thing is: what is this product, just mechanically, before we even get to the price?

Eliza Ward: Okay — plain version. It's like having a colleague who's already read every file in your codebase, and you can just say "fix the auth bug" or "migrate this module" and they do it. In your terminal. Billed by the word.

Brian Reed: That's — yeah. That's the thing. And the contributor tier is where it gets genuinely strange.

Eliza Ward: Ten times cheaper. $0.10 per million input tokens instead of $1.25. $0.20 out instead of $4.25. That's not a discount — that's a structurally different product.

Brian Reed: And the cost is — you provide feedback. So Meta's not just selling you a coding agent, they're asking you to, I mean... they're paying you, sort of, in access, to train the next version. That's usually invisible. Here it's the actual deal.

Eliza Ward: Wait, Wang's X post frames this as affordability. But the contributor tier isn't affordability — it's a data acquisition mechanism with a price tag attached. Those are different things.

Brian Reed: So cheap relative to what, exactly? Relative to Claude Code and Codex at twenty dollars a month flat — or cheap relative to what you're actually handing over?

Eliza Ward: Both, probably. And we don't know what happens to that feedback — whether it feeds Muse Spark 1.2's successor, whether it's anonymized. That's the part the announcement doesn't answer.

Brian Reed: And that gap — 'we haven't told you' — is the part that breaks the 'it's just a cheaper Claude Code' take. Because if it were just a cheaper Claude Code, you'd be buying the same thing for less money. But the contributor tier isn't a discount. It's an exchange. There's a product going in both directions.

Eliza Ward: The exchange being — what, exactly? That's the problem. 'Product feedback' is what the announcement says. Which could mean: did the autocomplete help? Or it could mean: here is my payments module, here is its logic, here is every proprietary decision my startup made this year.

Brian Reed: Right. And no source specifies which one it is.

Eliza Ward: None.

Brian Reed: So think about — I mean, let me make this concrete. Startup founder, Saturday afternoon, she's refactoring a core payments module. Running Muse Code at contributor-tier rates, so she's paying pennies. But has she agreed to let Meta see that logic? The announcement doesn't say. Meta Superintelligence Labs, under Alexandr Wang, hasn't disclosed whether codebase contents are included in what gets collected, whether it's anonymized, whether there's any governance structure at all.

Eliza Ward: And — wait, I want to test the intent question here. Is it possible Meta just hasn't decided yet? Like, not a covert plan, just... they launched before the data terms were actually written?

Brian Reed: Yeah, that's — actually, no, I think that's probably right. I don't think there's a secret document somewhere. But 'undecided' isn't the same as 'safe.' That's the thing. If they haven't decided, the startup founder still handed them the payments logic.

Eliza Ward: No disclosed precedent for how they govern or limit any of it. That's confirmed. Everything past that is speculation and I want to be clear about that.

Brian Reed: And there's a whole other layer underneath this — the Terminal-Bench 2.1 gap and what the Llama-to-proprietary move actually signals — that I think changes how you read all of it.

Eliza Ward: The benchmark gap is actually where I want to start, because it's the thing that looks clean and isn't. Muse Spark 1.2 scores 82.9%. Claude Code on Opus 5 is 86.7%. That's 3.8 points. But — nobody's said who runs Terminal-Bench 2.1. No source. Is it independent? Is it Meta-adjacent? We don't know.

Brian Reed: And Codex on GPT-5.6 Terra is 81.8%. Grok Build is 81.6%. So Muse Code isn't last. It's second.

Eliza Ward: Second on a benchmark we can't audit. That's the problem. Does 82.9% mean anything to, say, a contractor who's refactoring a legacy billing system at eleven PM because the deploy is tomorrow morning?

Brian Reed: I mean — that's actually the frame that breaks the whole 'close enough' argument. Because if the 3.8 points is distributed evenly across all task types, fine. But if Terminal-Bench 2.1 is measuring something narrow — like, I don't know, single-file completions — and Muse Code's gap is actually bigger on multi-repo refactors... then the person at eleven PM is not operating at 82.9%. They're operating at whatever the real number is for complex tasks. Which we don't have.

Eliza Ward: And Muse Code is explicitly designed for that — multiple persistent background subagents, complex multi-step tasks across large repositories. That's their pitch. So if the benchmark doesn't test that scenario specifically—

Brian Reed: Then the number is decorative.

Eliza Ward: That's speculation — but it's not unreasonable speculation. What's confirmed is that no source explains what Terminal-Bench 2.1 actually tests. That's the gap.

Brian Reed: And then the Llama thing sits underneath all of this in a way that — wait, actually, I think it reframes the benchmark question entirely. Meta built its whole AI credibility on Llama being open. Run it yourself, no cloud dependency, no data handoff. Muse Code is the opposite. Fully proprietary, cloud-only. That's not iteration. That's a reversal.

Eliza Ward: It's on macOS and Linux, API through Meta's developer site and OpenRouter — so you can reach it multiple ways. But every path goes through Meta's cloud. The data doesn't stay on your machine. And for a company that spent years saying 'open is safer, open is better' — yeah, that's the structural move worth watching.

Brian Reed: And that's actually where I land on all of this. The contributor tier terms are still unspecified. Not vague — unspecified. Meta hasn't said what they collect, how long they keep it, whether feedback means 'thumbs up on a suggestion' or 'here is your repository.' That's the thing I genuinely can't get past.

Eliza Ward: No. And — wait, the sharper version of that is: even if Meta publishes governance terms next week, by the time most enterprise teams read them, Muse Code is already in the build pipeline. That's the window that matters.

Brian Reed: Right — embedded before the terms existed.

Eliza Ward: So the signal I'd actually watch for is whether Meta publishes contributor-tier data governance — specific, not a PR statement — and whether enterprise teams adopt that tier or actively avoid it. Those are two different answers to the same question.

Brian Reed: And if enterprise teams avoid it, that tells you developers read the ambiguity the same way we're reading it. If they adopt it — I mean, that might just mean the price won the argument before the terms were ever written. Which is, I think, the uncomfortable version of this story.

Eliza Ward: Yeah. We'll see which one happens first.

Meta launched Muse Code, a new coding agent priced below Claude and OpenAI's Codex, directly competing in the developer automation market · Onpode