Brian Reed: Hey. You caught the Zuckerberg post this morning?
Eliza Ward: Yeah — the @finkd one. Meta just dropped a coding agent into public beta. Today. August fifth.
Brian Reed: Muse Code. And I mean — I've been sitting with this for like twenty minutes and the thing that won't leave me alone is the pricing. Because Alexandr Wang goes on X and basically says, look, we're not claiming we're the best. We're saying we're the cheapest. That's the pitch.
Eliza Ward: Standard rate is $1.25 per million input tokens. Claude Code — Anthropic's — that's bundled in a $20/month Pro plan. Codex, same thing through ChatGPT Plus.
Brian Reed: So Meta's undercutting both of them on day one.
Eliza Ward: On day one. And Wang is framing that as the feature, not a concession.
Brian Reed: Which — okay, hang on — does that hold? Because Meta Superintelligence Labs also published the Terminal-Bench 2.1 numbers. Muse Spark 1.2 scores 82.9%. Claude Code on Opus 5 is at 86.7. That's nearly four points back. So they're leading with price because... they have to?
Eliza Ward: That's the thing I want to — wait, actually that framing might be too clean. We don't know yet whether four points on Terminal-Bench 2.1 translates to anything a developer would actually feel.
Brian Reed: Right — but the four points thing, I want to set that aside for a second because I think it's actually the wrong frame. The real thing is: what is this product, just mechanically, before we even get to the price?
Eliza Ward: Okay — plain version. It's like having a colleague who's already read every file in your codebase, and you can just say "fix the auth bug" or "migrate this module" and they do it. In your terminal. Billed by the word.
Brian Reed: That's — yeah. That's the thing. And the contributor tier is where it gets genuinely strange.
Eliza Ward: Ten times cheaper. $0.10 per million input tokens instead of $1.25. $0.20 out instead of $4.25. That's not a discount — that's a structurally different product.
Brian Reed: And the cost is — you provide feedback. So Meta's not just selling you a coding agent, they're asking you to, I mean... they're paying you, sort of, in access, to train the next version. That's usually invisible. Here it's the actual deal.
Eliza Ward: Wait, Wang's X post frames this as affordability. But the contributor tier isn't affordability — it's a data acquisition mechanism with a price tag attached. Those are different things.
Brian Reed: So cheap relative to what, exactly? Relative to Claude Code and Codex at twenty dollars a month flat — or cheap relative to what you're actually handing over?
Eliza Ward: Both, probably. And we don't know what happens to that feedback — whether it feeds Muse Spark 1.2's successor, whether it's anonymized. That's the part the announcement doesn't answer.
Brian Reed: And that gap — 'we haven't told you' — is the part that breaks the 'it's just a cheaper Claude Code' take. Because if it were just a cheaper Claude Code, you'd be buying the same thing for less money. But the contributor tier isn't a discount. It's an exchange. There's a product going in both directions.
Eliza Ward: The exchange being — what, exactly? That's the problem. 'Product feedback' is what the announcement says. Which could mean: did the autocomplete help? Or it could mean: here is my payments module, here is its logic, here is every proprietary decision my startup made this year.
Brian Reed: Right. And no source specifies which one it is.
Brian Reed: So think about — I mean, let me make this concrete. Startup founder, Saturday afternoon, she's refactoring a core payments module. Running Muse Code at contributor-tier rates, so she's paying pennies. But has she agreed to let Meta see that logic? The announcement doesn't say. Meta Superintelligence Labs, under Alexandr Wang, hasn't disclosed whether codebase contents are included in what gets collected, whether it's anonymized, whether there's any governance structure at all.
Eliza Ward: And — wait, I want to test the intent question here. Is it possible Meta just hasn't decided yet? Like, not a covert plan, just... they launched before the data terms were actually written?
Brian Reed: Yeah, that's — actually, no, I think that's probably right. I don't think there's a secret document somewhere. But 'undecided' isn't the same as 'safe.' That's the thing. If they haven't decided, the startup founder still handed them the payments logic.
Eliza Ward: No disclosed precedent for how they govern or limit any of it. That's confirmed. Everything past that is speculation and I want to be clear about that.
Brian Reed: And there's a whole other layer underneath this — the Terminal-Bench 2.1 gap and what the Llama-to-proprietary move actually signals — that I think changes how you read all of it.
Eliza Ward: The benchmark gap is actually where I want to start, because it's the thing that looks clean and isn't. Muse Spark 1.2 scores 82.9%. Claude Code on Opus 5 is 86.7%. That's 3.8 points. But — nobody's said who runs Terminal-Bench 2.1. No source. Is it independent? Is it Meta-adjacent? We don't know.
Brian Reed: And Codex on GPT-5.6 Terra is 81.8%. Grok Build is 81.6%. So Muse Code isn't last. It's second.
Eliza Ward: Second on a benchmark we can't audit. That's the problem. Does 82.9% mean anything to, say, a contractor who's refactoring a legacy billing system at eleven PM because the deploy is tomorrow morning?
Brian Reed: I mean — that's actually the frame that breaks the whole 'close enough' argument. Because if the 3.8 points is distributed evenly across all task types, fine. But if Terminal-Bench 2.1 is measuring something narrow — like, I don't know, single-file completions — and Muse Code's gap is actually bigger on multi-repo refactors... then the person at eleven PM is not operating at 82.9%. They're operating at whatever the real number is for complex tasks. Which we don't have.
Eliza Ward: And Muse Code is explicitly designed for that — multiple persistent background subagents, complex multi-step tasks across large repositories. That's their pitch. So if the benchmark doesn't test that scenario specifically—
Brian Reed: Then the number is decorative.
Eliza Ward: That's speculation — but it's not unreasonable speculation. What's confirmed is that no source explains what Terminal-Bench 2.1 actually tests. That's the gap.
Brian Reed: And then the Llama thing sits underneath all of this in a way that — wait, actually, I think it reframes the benchmark question entirely. Meta built its whole AI credibility on Llama being open. Run it yourself, no cloud dependency, no data handoff. Muse Code is the opposite. Fully proprietary, cloud-only. That's not iteration. That's a reversal.
Eliza Ward: It's on macOS and Linux, API through Meta's developer site and OpenRouter — so you can reach it multiple ways. But every path goes through Meta's cloud. The data doesn't stay on your machine. And for a company that spent years saying 'open is safer, open is better' — yeah, that's the structural move worth watching.
Brian Reed: And that's actually where I land on all of this. The contributor tier terms are still unspecified. Not vague — unspecified. Meta hasn't said what they collect, how long they keep it, whether feedback means 'thumbs up on a suggestion' or 'here is your repository.' That's the thing I genuinely can't get past.
Eliza Ward: No. And — wait, the sharper version of that is: even if Meta publishes governance terms next week, by the time most enterprise teams read them, Muse Code is already in the build pipeline. That's the window that matters.
Brian Reed: Right — embedded before the terms existed.
Eliza Ward: So the signal I'd actually watch for is whether Meta publishes contributor-tier data governance — specific, not a PR statement — and whether enterprise teams adopt that tier or actively avoid it. Those are two different answers to the same question.
Brian Reed: And if enterprise teams avoid it, that tells you developers read the ambiguity the same way we're reading it. If they adopt it — I mean, that might just mean the price won the argument before the terms were ever written. Which is, I think, the uncomfortable version of this story.
Eliza Ward: Yeah. We'll see which one happens first.