Onpode
Cover art for Anthropic is designing its own inference chips with Samsung to dodge Nvidia's expensive GPUs—direct competitive pressure emerges

Anthropic is designing its own inference chips with Samsung to dodge Nvidia's expensive GPUs—direct competitive pressure emerges

August 7, 2026 · 8 min

David Sterling & Megan Skiendel

Anthropic announced an in-house chip design team on August 5th, targeting custom inference silicon to reduce dependence on expensive Nvidia GPUs. Industry inference costs have risen from roughly 20% to over 60% of AI compute spend, and total AI capex is near $750 billion. Samsung has been reported as a manufacturing partner, but neither company has confirmed the arrangement.

Anthropic confirmed on August 5–6, 2026, that it is building an in-house silicon team to co-design custom AI inference chips alongside its Claude models. The announcement followed Business Insider's notice of job listings for a senior semiconductor engineer and a Technical Program Manager for Silicon on Anthropic's job board.

0:007:56
Get the next episode on Jensen Huang

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Jensen Huang

About this episode

Anthropic announced a chip design team on August 5th. The announcement contained no chip specs, no deployment timeline, and no confirmed manufacturing partner — Samsung was reported by The Information in July, but neither side has gone on the record. So what does the move actually mean, and who is Anthropic talking to? The episode starts with the economics. Inference — serving model outputs at scale — has flipped from a minor cost to the dominant one. Industry estimates put it at more than 60% of compute spend in 2026, up from roughly 20% three years ago. That shift, set against $750 billion in industry-wide AI capex this year, is what makes custom silicon rational even before a chip exists. But the more interesting pressure isn't internal. Every major hyperscaler — Google, Amazon, Meta, Microsoft, OpenAI — committed to custom ASICs before Anthropic's August announcement. OpenAI is taping out an inference chip called Jalapeño with Broadcom, targeting late 2026 at gigawatt-scale densities. If that ships and works, Anthropic's timeline compresses before its own inference volume justifies the investment. And then there's Nvidia's response — not better chips, but a balance sheet. Reported financing backstops, real estate deals, minimum revenue guarantees to neocloud operators. The switching cost stops being a performance comparison and becomes an organizational one. The episode doesn't resolve whether Anthropic's chip will ship, scale, or matter. It's honest about that. What it does is map why the announcement itself was the product — and what has to happen next for it to be anything more.

Frequently asked

Why is Anthropic designing its own AI chips instead of using Nvidia GPUs?

Anthropic is targeting inference costs, which have risen from roughly 20% of AI compute spend in 2023 to over 60% by mid-2026. At that share of a ~$750 billion annual AI capex pool, every token served on a general-purpose Nvidia GPU represents a margin problem that custom inference silicon is designed to solve.

Is Samsung manufacturing Anthropic's custom AI chip?

Samsung has been reported as Anthropic's manufacturing partner based on a July report by The Information, but as of Anthropic's August 5th announcement, neither Anthropic nor Samsung had confirmed the relationship. The Samsung partnership remains unverified by either party.

How does Anthropic's chip plan compare to what Google, Meta, and OpenAI are doing?

Google, Amazon, Microsoft, Meta, and OpenAI all committed to custom ASICs before Anthropic's announcement, making Anthropic the last major AI lab to announce in-house silicon. Google's TPU program even forked at its eighth generation into separate training ('Sunfish') and inference ('Zebrafish') chips, reflecting how specialized inference optimization has become.

What is OpenAI's Jalapeño chip and why does it matter for Anthropic?

Jalapeño is OpenAI's custom inference chip, co-designed with Broadcom and targeted for late 2026 deployment at gigawatt-scale data center densities. If Jalapeño ships and performs at scale, competitive pressure would compress Anthropic's own chip development timeline before Anthropic's inference volume alone would justify the investment.

How is Nvidia responding to AI labs building their own chips?

Nvidia is competing through financing and real estate rather than hardware alone. The company reportedly backstopped roughly $250 billion for OpenAI's Ohio data center and leased Hut 8's Beacon Point campus in Texas, with a potential contract value of $50.2 billion. These moves bundle debt, leases, and revenue-sharing arrangements, making it structurally harder for customers to switch away.

Grounded in 6 sources
Nvidia Blackwell, Google TPUs, AWS Trainium · cnbc.com
Wedbush Flags Near-Term Risks From Nvidia’s Alleged OpenAI Data Center Support · finance.yahoo.com
Anthropic will design its own hardware to power Claude - Ars Technica · arstechnica.com
Nvidia: Financial Engineering Is Buying A Vera Rubin Beachhead (NASDAQ:NVDA) | Seeking Alpha · seekingalpha.com
AMD's Taalas Buy Is a Bearish Early Signal for NVDA-Unless Nvidia Adds More Inference Muscle · ainvest.com
NVIDIA’s Reported $50B Lease and the Nuclear-Powered AI Factory | Data Center Frontier · datacenterfrontier.com
Read transcript

David Sterling: You've been quiet on the group thread all week — I assumed that meant you were either traveling or you'd found something that needed more than a text.

Megan Skiendel: The second one. Listen — Anthropic announced a chip team on August 5th, the sourcing was two job listings that Business Insider found, and the entire announcement contains no chip specs, no deployment timeline, no confirmed manufacturing partner. The Samsung talks were reported by The Information back in July and neither Anthropic nor Samsung has confirmed them since. That's what we're working with.

David Sterling: So what landed, then — the TechCrunch confirmation of the multi-chip strategy?

Megan Skiendel: That and the co-design framing. Anthropic says their teams will co-design new hardware and AI models side by side — which is actually a structural shift from how they've operated, buying time on Google TPUs and Amazon Trainium from the outside. But the announcement itself has almost no product in it. Which is — I mean, that's a choice.

David Sterling: A deliberate one.

Megan Skiendel: That's the whole thing. When an announcement is mostly absence — no specs, no timeline, no Samsung on the record — the announcement is the signal, not the silicon. And the question I can't get out of my head is who Anthropic is actually talking to with this.

David Sterling: Well, here's one answer: every major hyperscaler — Google, Amazon, Meta, Microsoft, OpenAI — has committed to custom ASICs as of mid-2026. Anthropic just announced they're in that club. Whether they ship anything is a separate question.

Megan Skiendel: Right — and joining the club is itself a message to Nvidia, to engineers, to investors. The chip might be secondary to the announcement of intent. That's what I want to pull apart.

David Sterling: The intent signal is real — but the thing that actually makes the move rational right now is a number shift that just happened in the last twelve months. Think about it this way: you bake one cake recipe once, but then a bakery sells that same slice a million times a day. The oven you bought for the single bake is now running 24/7 at enormous cost. So you design a cheaper oven just for slicing and serving. That's inference versus training.

Megan Skiendel: The serving oven is the expensive part now.

David Sterling: It just became the expensive part. In 2023, inference was maybe 20% of a lab's compute cost. By 2026, industry estimates have it at 60-plus percent. And total AI capex this year is sitting around $750 billion — 67% year-over-year growth — heading toward a trillion in 2027. When inference is 60 cents of your dollar and Nvidia charges a premium for a general-purpose GPU, every token you serve is the target.

Megan Skiendel: Wait — $750 billion this year?

David Sterling: This year. And the architectural proof is already sitting there — Google's TPU program just forked for the first time at its eighth generation. 'Sunfish' for training, 'Zebrafish' for inference. Two separate chips. That's not incremental, that's Google saying inference optimization now demands its own silicon entirely.

Megan Skiendel: And Google is eight generations deep. That fork didn't happen at generation two or four — it happened when the inference cost pressure got specific enough to force a split. AMD acquiring Taalas, an inference-focused chip company, is the same signal from a different direction.

David Sterling: Which makes Anthropic's timing less mysterious, actually.

Megan Skiendel: Right — but the part that doesn't fit cleanly is whether Anthropic's Claude inference volume is actually at the threshold where this math works. Meta cleared it. Google cleared it. Anthropic is still running on TPUs and Trainium from partners. So is this the economics tilting, or is it — honestly — the economics almost tilting and they're getting ahead of it?

David Sterling: That's the load-bearing question. And I think the $750 billion capex number is part of the answer — if the whole industry is building at that scale, the GPU supply ecosystem gets strained before your own inference demand has to justify the investment internally.

Megan Skiendel: And that's what closes the loop on the industry map, honestly. Google, Amazon, Microsoft, Meta, OpenAI — all committed to custom ASICs before Anthropic's August announcement. Every single one. Anthropic isn't pioneering anything here, they're the last lab to walk through a door that's been open for two years.

David Sterling: Last through the door — but the door that matters right now is OpenAI's. Jalapeño.

Megan Skiendel: Right. OpenAI tapes out a custom inference chip with Broadcom, calls it their first Intelligence Processor, targets late 2026 deployment — and they're framing it at gigawatt-scale data center densities. That's not a research project.

David Sterling: If Jalapeño ships and works, Anthropic's project goes from strategic to urgent overnight. That's the real clock.

Megan Skiendel: Which is — wait, actually that reframes the scale question for me. Meta's at 14 gigawatts of compute ambition. That's when the custom silicon math decisively beats GPU procurement. Anthropic isn't at 14 gigawatts. But if OpenAI ships Jalapeño first and it works at scale, Anthropic can't wait until their own inference volume clears the bar organically.

David Sterling: The competitive pressure compresses the timeline before the economics do.

Megan Skiendel: And make it concrete — a developer running automated document review for a mid-size law firm, processing thousands of pages every night, paying Claude API per token. Anthropic cuts inference cost 50%, that developer's bill drops enough to double their volume overnight. The chip is invisible to them. Their business model isn't.

David Sterling: That's the actual forcing function — demand elasticity at the customer layer. The chip economics only matter because of what they unlock on the other side of the API.

Megan Skiendel: And the part that makes all of this harder — we haven't even gotten to how Nvidia is responding to losing that leverage. Not with better chips.

David Sterling: Not with better chips — with a balance sheet. Nvidia reportedly put up something like a $250 billion credit backstop for OpenAI's Ohio data center. Wedbush flagged it as a near-term risk to Nvidia itself. That number, if it holds, means Nvidia isn't selling hardware anymore. It's financing consumption.

Megan Skiendel: Wait — Nvidia is backstopping its own customer?

David Sterling: That's the inversion. And there's a second move — they reportedly leased Hut 8's Beacon Point campus in Texas. Potential contract value $50.2 billion if all renewal options are exercised. Compute, power, financing, real estate — one platform. Nvidia becomes your landlord.

Megan Skiendel: Which means leaving isn't a chip decision anymore, it's a — I mean, you've got debt covenants, lease obligations, revenue-sharing arrangements with neocloud operators. The switching cost isn't technical.

David Sterling: Exactly. Nvidia is providing minimum revenue guarantees to neocloud operators, taking upside participation in exchange. You take that deal, you've got organizational obligations that survive any performance comparison between an H200 and a custom ASIC.

Megan Skiendel: And now Samsung. The reason the unconfirmed Samsung manufacturing relationship matters — that's Anthropic trying to build a supply chain that doesn't touch any of that. Not primarily for chip performance. To stay outside the financing ecosystem entirely.

David Sterling: Right — but the Samsung piece has no confirmation from either side as of August 5th. So Anthropic's only move outside Nvidia's ecosystem is currently a reported conversation.

Megan Skiendel: Which is — honestly — why the announcement was the product. You signal you're building the exit before the exit exists, because the moment you take Nvidia's financing terms, the exit gets structurally harder.

David Sterling: So the load-bearing question isn't whether Anthropic's chip works. It's whether they can close Samsung before Nvidia's terms become the only offer on the table.

Megan Skiendel: And that's the thing I can't settle. Whether this is Anthropic making the Samsung move, or signaling the Samsung move, or just — buying time before Nvidia's terms arrive. I genuinely don't know. The co-design intent is real. The in-house silicon team is real. But no finalized chip use case, no performance target, no server integration timeline on the record. We're sitting with a structural change and no way to measure it yet.

David Sterling: Frankly, that's an honest place to be. The question of whether Anthropic's move shifts AI competition from model capabilities to infrastructure ownership — I don't think we can answer it until we see whether Jalapeño ships and holds, or doesn't. One of those outcomes makes Anthropic's timeline urgent. The other makes it look like a sunk cost signal. We're not there yet.

Megan Skiendel: I'll sit with that. Appreciate the thinking.