Onpode
Cover art for Nvidia's Vera Rubin combines CPUs and GPUs into one platform—a direct challenge to AMD and Intel's hold on the CPU side

Nvidia's Vera Rubin combines CPUs and GPUs into one platform—a direct challenge to AMD and Intel's hold on the CPU side

July 22, 2026 · 9 min

Ryan Castillo & Jordan Hale

Nvidia's Vera Rubin NVL72 — 72 Rubin GPUs and 36 Vera CPUs fused via NVLink-C2C into one rack-scale system — claims 5x the inference performance of Blackwell at 10x lower cost per token. But Nvidia never published AMD EPYC 9755's raw benchmark scores, and Anthropic's 2-gigawatt AMD commitment signals real lock-in anxiety.

Nvidia's Vera Rubin platform, unveiled at GTC 2025 and detailed further at SIGGRAPH 2026 and ISC High Performance 2026, represents a rack-scale AI supercomputer that tightly integrates CPUs, GPUs, networking, and storage into a single coherent system.

0:008:57
Make your own on Onpode

Describe any topic. Hear it in minutes.

More Onpode episodes on Nvidia

About this episode

Nvidia's Vera Rubin platform arrived with the kind of coverage that mostly marvels and moves on. This episode tries to slow that down. The hardware is genuinely striking — 336 billion transistors in the Rubin R100 GPU, an 88-core Olympus CPU architecture that's Nvidia's first, and a chip-to-chip interconnect called NVLink-C2C that eliminates the memory boundary between CPU and GPU entirely. That last part is the real news, and it's also what makes the competitive question so thorny: once you've welded the CPU and GPU together into a rack-scale system of 72 GPUs, 36 CPUs, and five co-designed tray components, you're not buying a chip anymore. You're buying a system, or you're buying nothing. The episode examines what that means for AMD and Intel, who can compete on a core but cannot replicate a 350-factory, 30-country supply chain on any timeline that matters. It also looks hard at the benchmark Nvidia used to claim Vera edges out AMD's EPYC 9755 — and the fact that AMD's raw scores were never published, only Nvidia's aggregate. That opacity turns out to have consequences: it's part of what's driving Anthropic's commitment to purchase up to 2 gigawatts of AMD silicon, a number too large to call a backup plan. The episode ends somewhere genuinely uncomfortable — because nobody's answered yet whether AMD's counter-bet lands before Nvidia's manufacturing web makes the exit irrelevant.

Frequently asked

What is Nvidia Vera Rubin and what makes it different from previous GPU platforms?

Nvidia Vera Rubin is a rack-scale AI system combining 72 Rubin GPUs, 36 Vera CPUs, ConnectX-9 SuperNICs, and BlueField-4 DPUs — seven co-designed chips — into one integrated platform. Its core innovation is NVLink-C2C, which eliminates the memory boundary between CPU and GPU so they share the same address space directly.

How does Nvidia Vera Rubin performance compare to Blackwell?

Nvidia claims the Vera Rubin NVL72 delivers 5x the inference performance of Blackwell at 10x lower cost per token. Nvidia attributes the gain primarily to NVLink-C2C integration — which collapses CPU-GPU latency — and the Rubin R100 GPU's 336 billion transistors. No independent source has verified these figures.

Did Nvidia publish AMD EPYC benchmark scores when comparing the Vera CPU?

Nvidia did not publish AMD EPYC 9755 raw scores when announcing the Vera CPU. Nvidia controlled the test environment, benchmark selection, and released only an aggregate comparison. Neither AMD nor any independent source has had the opportunity to challenge the methodology or verify the per-core performance claims.

Why is Anthropic buying $5 billion worth of AMD chips if Nvidia Vera Rubin is so powerful?

Anthropic committed to purchase up to 2 gigawatts of AMD silicon — a scale that signals a deliberate infrastructure counter-bet, not a hedge. Because Nvidia never released EPYC's raw benchmark scores, hyperscalers like Anthropic cannot independently verify Nvidia's performance claims and are building an exit route before rack-scale lock-in is complete.

Does Nvidia Vera Rubin create vendor lock-in?

Nvidia Vera Rubin's NVL72 bundles seven co-designed chips across five rack-scale tray systems — buyers cannot purchase the GPU layer from Nvidia and the CPU layer from AMD or Intel separately. Nvidia's 350 manufacturing sites across 30 countries deepens the switching cost beyond architecture into supply-chain infrastructure that AMD and Intel cannot replicate quickly.

Grounded in 12 sources
Exclusive: Nvidia's Jensen Huang defends Chinese AI amid Kimi panic - Axios · axios.com
Wistron shares surge in Taiwan after opening of Texas AI superchip plant to supply Nvidia - CNBC · cnbc.com
NVIDIA ships Vera CPUs to OpenAI, Anthropic and SpaceX · cnbc.com
AMD to invest up to $5 billion in Anthropic; AI startup to buy up to 2 GW of chips - Reuters · reuters.com
Choose Your Fighter: Nvidia CEO and Jim Cramer Offer Dueling Visions of AI’s Future - Gizmodo · gizmodo.com
Nvidia Wants to Own Every Chip Inside AI Data Centers · wired.com
Nvidia details AI rack strategy with Vera Rubin NVL72 and custom Vera CPU · digitimes.com
NVIDIA AI Infrastructure posts Vera performance claims · hothardware.com
Nvidia Unveils Vera AI CPU, Taking Aim at AMD And Intel In The Battle For Data Centers | IBTimes · ibtimes.com
NVIDIA ramps up Vera Rubin AI system for cloud giants · itbrief.news
Next Gen Data Center CPU | NVIDIA Vera CPU · nvidia.com
Nvidia Vera Rubin: Inside the agentic AI factory that rewrites the CPU playbook - SiliconANGLE · siliconangle.com
Read transcript

Jordan Hale: Ryan, hey — I want to start somewhere a little uncomfortable, if that's okay.

Ryan Castillo: That's usually where the good stuff is — go.

Jordan Hale: So Nvidia just launched Vera Rubin, and the coverage has been, like... glowing. Almost uniformly. And I think part of what we want to do today is slow that down, because when I actually look at what was announced — the NVL72, which is 72 Rubin GPUs and 36 Vera CPUs in one rack-scale system, already at CoreWeave and Google Cloud and Azure and Oracle — my first reaction was awe. My second reaction was: who exactly is this story being told for?

Ryan Castillo: What do you mean by that?

Jordan Hale: I mean Jensen Huang is a brilliant narrator. And when you roll out a CPU — an 88-core 'Olympus' architecture, Nvidia's first — and you bench it against AMD's EPYC 9755 but you don't release AMD's actual scores, only your aggregate... you know, that's a very careful story.

Ryan Castillo: That's not a small thing to flag — withholding the opponent's raw numbers while claiming a win is the kind of move that should make any analyst pause.

Jordan Hale: And yet the Rubin R100 GPU has 336 billion transistors. Like — the hardware might genuinely be that good. That's what makes this hard.

Ryan Castillo: Right — but the question for the whole episode is whether 'genuinely good hardware' and 'vendor lock-in at rack scale' are two separate things or the same thing dressed up. And Intel and AMD are about to find out the hard way which it is.

Jordan Hale: But that 'same thing dressed up' framing — I think that's actually where we need to slow down, because the technical reason this is different from every previous GPU upgrade is kind of... buried in the announcement.

Ryan Castillo: NVLink-C2C. That's the actual news.

Jordan Hale: Okay — explain it like I'm not a chip person.

Ryan Castillo: Your laptop has a CPU and a GPU that pass notes to each other through a slow mailbox. Vera Rubin tears out the mailbox and welds the two chips together so they share the same memory directly. That's it. That's what NVLink-C2C does — it collapses the address space boundary between the Vera CPU and the Rubin GPU entirely. They're not coordinating anymore. They're one thing.

Jordan Hale: Wait — so when Nvidia claims 5x the inference performance of Blackwell at 10x lower cost per token, that gap isn't just more transistors. It's... the elimination of that latency.

Ryan Castillo: That's the load-bearing piece, yeah. And the reason it matters structurally — not just technically — is that once you've done that, you can't just buy the GPU layer from Nvidia and the CPU layer from AMD or Intel anymore. The NVL72 bundles 72 Rubin GPUs, 36 Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs — seven co-designed chips, five rack-scale tray systems. You're buying a system or you're buying nothing.

Jordan Hale: And here's the number that actually stopped me — 350 factory sites. Across 30 countries. For one platform. I mean, that's not a chip company's supply chain. That's... you know, that's how you build a car.

Ryan Castillo: That's the systems company signal. AMD and Intel don't have that industrial footprint. They can compete on a core, maybe a socket. They cannot replicate that supply chain in any timeframe that matters for the next procurement cycle.

Jordan Hale: Which is exactly why the AMD-Anthropic move is so interesting — because if the integration is this real, why is Anthropic committing to 2 gigawatts of AMD silicon? That's not a hedge. That's a counter-bet.

Ryan Castillo: That counter-bet framing is right — but it assumes Anthropic actually verified what they're buying performs the way Nvidia says it does. And that's where the circulating take falls apart. Everyone's reporting Vera beat AMD's EPYC 9755. That's not what the data shows. What Nvidia released was an aggregate comparison. EPYC's raw scores never appeared.

Jordan Hale: Wait — they literally didn't publish the other side's numbers?

Ryan Castillo: Correct. Nvidia controlled the test environment, the benchmark selection, and what got published. You got Nvidia's aggregate. You did not get AMD's score.

Jordan Hale: Okay but — I mean, that's not just a transparency issue, right? Like, when you control the test and the narrative, you don't actually need to win on the numbers. You just need to win on the vibe. And the vibe is, Vera CPU edges out EPYC 9755 in per-core performance under fully loaded socket conditions. Nobody's going back to check that.

Ryan Castillo: No independent source has verified it. Not for general-purpose data center workloads — only for the inference-specific conditions Nvidia designed the test around.

Jordan Hale: Which, wait — that's actually... the benchmark conditions may not reflect what hyperscalers actually run on CPUs day to day. This thing is framed as purpose-built for agentic AI, but most CPU workloads in a data center aren't agentic inference. So even if Vera wins on Nvidia's test, we don't know what it does on anything else.

Ryan Castillo: Look — even structurally opaque benchmarks signal something. The fact Nvidia chose EPYC 9755 as the comparison and still wouldn't release the raw scores suggests the margin is narrow enough to control carefully. You don't bury numbers you're dominating by.

Jordan Hale: And Intel hasn't responded to these specific claims at all. AMD either. Neither incumbent has had a chance to challenge the methodology.

Ryan Castillo: Which connects directly to why hyperscalers are hedging — and we'll get into how AMD's five billion dollar Anthropic investment is partly a response to not being able to verify these performance claims independently. That's the thread that makes this benchmark opacity genuinely consequential.

Jordan Hale: And that's the part that makes the AMD five billion dollar Anthropic move feel less like a hedge and more like — I mean, like a fire alarm someone pulled early. Because Anthropic committing to purchase up to 2 gigawatts of AMD chips, that's not a backup plan. That's infrastructure at a scale that only makes sense if you genuinely believe Nvidia's performance story can't be independently verified and you want your own escape route before it matters.

Ryan Castillo: The number that matters there is the 2 GW commitment — that's the tell. You don't contract 2 gigawatts of capacity for a workload you're not planning to run at scale. That's counter-infrastructure, not insurance.

Jordan Hale: Right — but the part that doesn't fit is that Anthropic is also a Vera Rubin customer, right? Google Cloud, Azure, Oracle — these same organizations are the ones deploying NVL72 racks *and* funding the alternative simultaneously.

Ryan Castillo: That's the structural trap Nvidia built for itself. They need hyperscaler adoption to validate the platform — but every hyperscaler that adopts at scale is now the customer with the most incentive to fund an exit before lock-in is complete.

Jordan Hale: So picture an infrastructure procurement lead — say it's late 2026, she's reviewing a CoreWeave quote for Vera Rubin NVL72 racks at integrated system pricing. And sitting next to that quote is AMD's Anthropic commitment. She's not choosing chips. She's choosing who she negotiates with in 2029.

Ryan Castillo: And she can't even validate the performance delta that justifies the CoreWeave price — because Nvidia never released EPYC 9755's raw scores.

Jordan Hale: Which, you know, that's the loop closing. The opacity creates the hedge, which funds the competition, which is the thing Nvidia's integration is supposed to prevent.

Ryan Castillo: Watch Wistron. Their Taiwan shares spiked after they opened the Texas AI superchip plant for Nvidia's supply chain. That's the signal the entrenchment is already financial — it's not just architectural anymore. Once that manufacturing web is load-bearing, the switching cost isn't just a chip decision.

Jordan Hale: So the thing to actually watch isn't whether Vera CPU beats EPYC 9755 on some future independent benchmark — it's whether AMD can make that 2 GW commitment real before Wistron's Texas plant makes the alternative supply chain irrelevant.

Ryan Castillo: The watch signal for me — if hyperscalers accelerate their custom silicon timelines in the next twelve months, that's the tell. Not that Nvidia lost. That the lock-in worked too well and they're running before the door closes.

Jordan Hale: And that's the question I can't shake, you know. Like — does Vera Rubin's integration make defection faster or impossible? Because AMD's $5 billion Anthropic bet is the clearest signal that Nvidia's largest customers believe the lock-in is real. But I don't actually know if believing it's real means you can escape it, or... I mean, that's the thing nobody's answered yet.

Ryan Castillo: No. And I don't think we can answer it today.

Jordan Hale: That's an uncomfortable place to stop.

Ryan Castillo: Yeah. Good conversation, though — genuinely.

Nvidia's Vera Rubin combines CPUs and GPUs into one platform—a direct challenge to AMD and Intel's hold on the CPU side · Onpode