Onpode
Cover art for Nvidia is moving away from selling chips individually—now it's all about AI racks

Nvidia is moving away from selling chips individually—now it's all about AI racks

August 1, 2026 · 10 min

Max Rivera & Clara Bennett

Nvidia's GB300 NVL72 rack delivers 1,648 teraflops per GPU on DeepSeek-V3 pre-training — a benchmark only possible because NVLink and 130 TB/s bandwidth are co-designed into the rack. Meanwhile, custom ASICs are growing 44.6% year-over-year versus 16.1% for GPUs, signaling inference volume is already moving away from Nvidia.

Nvidia is executing a significant strategic pivot away from selling individual GPUs toward delivering complete, integrated AI rack systems — sometimes called "AI factories" — that bundle processors, networking, cooling, and software into turnkey deployments.

0:009:58
Get the next episode on Nvidia

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Nvidia

About this episode

When Alibaba's stock moved five percent on a rumored compute deal for Moonshot AI, markets were pricing a transaction whose actual contents nobody had confirmed. That ambiguity turns out to be a useful door into something bigger: Nvidia has quietly changed what it sells. The unit is no longer a chip — it's an AI factory. Full rack, pre-wired, pre-plumbed, software included. One deployable infrastructure platform. The GB300 NVL72's benchmark on DeepSeek-V3 pre-training is real, and it's only achievable because the networking, cooling, and compute are co-designed from the ground up. Foxconn and Quanta, once system designers, are now assemblers of Nvidia's reference architecture. This episode follows the strategic logic of that shift — and then presses on its limits. Custom ASICs grew 44.6% year-over-year against GPU growth of 16.1%. Inference, not pre-training, is where AI compute volume actually runs — and that's precisely where Broadcom and AMD have a meaningful cost advantage. The rack-scale pivot is partly offensive, partly defensive cover. Whether the lock-in compounds or fractures depends on a question that isn't settled: do hyperscalers get paid back on AI infrastructure investment? If they do, Nvidia's model gets validated. If they don't, the switching cost starts pointing the wrong direction. Neither outcome is priced yet.

Frequently asked

What is the Nvidia GB300 NVL72 and what performance does it achieve?

The Nvidia GB300 NVL72 is a rack-scale AI system housing 72 Blackwell GPUs connected via fifth-generation NVLink, delivering 130 terabytes per second all-to-all bandwidth. On July 21st it set a benchmark of 1,648 teraflops per GPU on DeepSeek-V3 pre-training — a result Nvidia says cannot be replicated with individually assembled components.

Why is Nvidia selling complete AI racks instead of individual GPUs?

Nvidia shifted to rack-scale systems because peak AI performance now requires co-designed chips, networking, and cooling that cannot be matched by assembled parts. Nvidia CFO Colette Kress frames returns to investors at the infrastructure level, not per chip — a framing shift that Bank of America argues the market has not yet fully priced.

Are custom AI chips like Google TPUs and Amazon Trainium threatening Nvidia?

Custom ASICs — including Google TPUs, Amazon Trainium, and Broadcom designs — grew 44.6% year-over-year versus 16.1% for GPUs, nearly three times faster. The cost advantage is sharpest in inference workloads, where custom silicon runs 40–65% cheaper than general-purpose GPUs, which is precisely where AI compute volume is highest.

What does the Nvidia rack-scale strategy mean for manufacturers like Foxconn and Quanta?

Under Nvidia's Level-10 systems approach, starting with the Vera Rubin platform, Nvidia supplies fully tested server trays and partners like Foxconn, Quanta, and Wistron assemble them. This reduces those manufacturers from system designers to system integrators — they bolt Nvidia's design together but no longer own the design.

Is Nvidia's AI infrastructure lock-in as strong as investors believe?

Nvidia's rack-scale lock-in is real but selective. Neoclouds like CoreWeave lack the engineering depth to switch and remain fully committed to Nvidia GPU fleets. However, hyperscalers with their own silicon programs — Google, Amazon — have genuine optionality. Bank of America argues the market still prices Nvidia as a chip vendor, understating both the switching cost and the concentration risk.

Grounded in 7 sources
Nvidia shows no sign of AI slowdown after data center soars over 400% · cnbc.com
Alibaba shares rise on Moonshot’s 20,000-chip AI deal · finance.yahoo.com
Nvidia slips 4% as blockbuster AI spending spree triggers balance sheet jitters - Yahoo Finance · finance.yahoo.com
China’s tech advances are causing chaos from Silicon Valley to the White House - The Guardian · theguardian.com
Disaggregated Inference Is Splitting AI Hardware In Two - Forbes · forbes.com
Nvidia's CEO says ‘a lot’ of six-figure jobs in plumbing and construction are about to be unlocked - Fortune · fortune.com
Moonshot trained Kimi K3 on 20,000 Nvidia chips from Alibaba’s cloud · thenextweb.com
Read transcript

Clara Bennett: Max, I want to start with a question I couldn't answer this week — when Alibaba's stock moves five percent on a compute deal for Moonshot AI, what exactly did the market think it was buying?

Max Rivera: Wait, that's — yeah, that's the exact thing. Because the deal is 20,000 chips announced as a single arrangement, reported July 31st, and Alibaba came out and said, we did not supply H200 processors. Period. Did not say: we didn't do the deal.

Clara Bennett: So the market is pricing a transaction whose actual contents nobody has confirmed.

Max Rivera: Which is — I mean, that's not an accident. Jensen Huang has been deliberate about this. Nvidia stopped announcing GPU counts. They announce clusters now. The unit of sale is infrastructure, not silicon. And when you look at what that's built into — $39.1 billion, single quarter, Data Center revenue, 89% of the whole company — that is not a chip business anymore by any normal definition.

Clara Bennett: Now the question is whether that concentration is the moat or the fragility.

Max Rivera: That's it. That's the whole episode. Because Moonshot AI and its Kimi models are sitting on a rack of unknown silicon through Alibaba, markets are moving, and nobody's asking what's actually in the box.

Clara Bennett: But what's actually in the box — that's the wrong question to start with. The architecture is the confession, not the opacity. Think of it this way: Nvidia stopped selling kitchen appliances. They're selling the whole commercial kitchen — pre-wired, pre-plumbed, pre-staffed. That's the DSX AI factory. Chips, networking, cooling, software — one deployable unit.

Max Rivera: Wait — so the Moonshot deal being opaque is almost beside the point?

Clara Bennett: In practice, yes. Because the GB300 NVL72's 1,648 teraflops per GPU on DeepSeek-V3 pre-training — that benchmark, July 21st — you cannot hit that number with assembled parts. It requires fifth-generation NVLink and 130 terabytes per second all-to-all bandwidth co-designed into the rack. The rack is the product.

Max Rivera: Okay, but — wait, that's actually the part I keep tripping on. If Nvidia is designing the full rack, what does that make Foxconn? Or Quanta, Wistron?

Clara Bennett: Assemblers. That's the Level-10 systems move — Nvidia supplies fully tested server trays starting with the Vera Rubin platform, and Foxconn, Quanta, Wistron go from system designers to system integrators. They bolt Nvidia's design together. They don't own the design anymore.

Max Rivera: That's — yeah, that's a meaningful demotion.

Clara Bennett: And Bank of America's argument is that the market hasn't caught up to this. They're still pricing Nvidia as the dominant accelerator vendor — a chip company — while the actual strategic objective has shifted to systems platform. The switching cost isn't just CUDA anymore. It's the entire operational architecture of an AI factory.

Max Rivera: So — I mean, is that framing solid though? Because hyperscalers like Google, they're not buying commercial kitchens. They want to build their own. The custom ASIC growth, 44.6% year-over-year against GPU growth at 16.1% — that gap exists precisely because some customers are opting out.

Clara Bennett: That's the real tension. The rack-scale model is structurally strong on pre-training — the GB300 number proves that. But most AI spend by volume is inference, and that's exactly where the custom silicon cost edge, Broadcom and AMD, is sharpest. Nvidia's pivot is sound. Whether it holds at the inference layer — that's genuinely unsettled.

Max Rivera: Right — and inference being the fault line, that's not theoretical. I mean, picture a radiologist running image diagnostics in Nairobi. Thousands of inference queries per hour. That workload doesn't need the GB300 NVL72's 1,648 teraflops on pre-training — it needs cheap, fast, repetitive inference, which is exactly what Broadcom's custom silicon is built for. And at 40 to 65 percent lower cost than a general-purpose GPU, the math just — it doesn't work in Nvidia's favor there.

Clara Bennett: That cost gap is real. And importantly, that's not the niche — inference is where the volume runs.

Max Rivera: Which is why the 44.6 versus 16.1 number is the tell. Custom ASICs — Google TPUs, Amazon Trainium, Broadcom — growing nearly three times faster than GPUs. That's not noise.

Clara Bennett: No, it isn't. And notice how Colette Kress framed the return story to investors — cloud providers seeing 'immediate and strong return' on AI chip investment, but she said it at the infrastructure level. Not per chip.

Max Rivera: Wait — you're saying Nvidia's own CFO is already not telling the per-chip story?

Clara Bennett: In practice, yes. When the per-chip story gets harder to tell, you move the frame up. The rack-scale pivot isn't just offensive positioning — it's also cover.

Max Rivera: So — okay, that's the partial win for the hot take, right? The rack move is genuinely partly defensive. Nvidia sees custom silicon eating inference and moves upmarket to pre-training and mixed-precision workloads before the floor drops out. The GB300 benchmark is real, it's enormous, but it's benchmarking the workload Nvidia still owns — not the one growing fastest.

Clara Bennett: That's the kernel. Though the part we haven't touched yet — whether the lock-in actually compounds or whether hyperscalers have enough leverage to fracture it — that changes the valuation math completely, and neither outcome is priced in right now.

Max Rivera: Yeah — and honestly, I'm not sure I know which way that breaks. Yet.

Clara Bennett: The fracture point, though — it's not hyperscalers in general. It's specifically the ones with enough engineering depth to absorb the switching cost. CoreWeave is scaling hard on Nvidia GPU fleets right now. That's the platform adoption case in real time. They're not building their own silicon. They're betting the rack-scale model holds.

Max Rivera: Wait — CoreWeave is kind of the best example of who does NOT have the leverage to opt out, right? They're a neocloud. They don't have Google's TPU program or Amazon Trainium. So of course they're all-in on Nvidia.

Clara Bennett: Exactly — and that distinction matters. The CUDA ecosystem plus the full rack integration, the switching cost isn't just 'we'd have to retrain our models.' It spans hardware, software, and the operational know-how of running an AI factory. CoreWeave can't unbundle that on a reasonable timeline. Alibaba, maybe. But Alibaba is also the company that obscured what silicon is in the Moonshot cluster.

Max Rivera: Which — yeah, that opacity cuts both ways. If Alibaba had built a clean alternative, why not say so?

Clara Bennett: Right, and AMD is trying to compete at the rack level — so it's not that no one's attempting it. But 'attempting' and 'replacing an AI factory architecture' are very different engineering projects. Bank of America's point is that the market prices Nvidia like a chip vendor when the actual switching cost is closer to replacing your entire infrastructure team's institutional knowledge.

Max Rivera: That's — I mean, that's a harder number to put in a model. Operational know-how doesn't show up on a spec sheet.

Clara Bennett: It doesn't. So here's the calibrated version: the lock-in is real and it compounds — but only below a certain engineering threshold. Above that threshold, the Alibabas of the world have genuine optionality. The unresolved question is whether hyperscaler returns on AI infrastructure actually materialize. If they do, Nvidia's multiple re-rates upward — the rack model gets validated. If those returns disappoint, the lock-in starts looking like leverage pointed the wrong direction.

Max Rivera: So the valuation bet is really a bet on whether the people writing the biggest checks get paid back first.

Clara Bennett: That's the whole thing. And neither outcome is priced yet. Which means Nvidia is simultaneously the most defensible infrastructure company in the world — and the one with the most undisclosed risk sitting in that exact question.

Max Rivera: So Nvidia might be selling complete kitchens — I'll give you that — but if the hyperscalers decide they'd rather be Michelin-star chefs than restaurant owners, those kitchens become very expensive furniture. The GB200 and GB300 NVL72 are genuinely impressive infrastructure, I'm not disputing that. But 44.6% versus 16.1% — that's the number that matters to me. Not the benchmark records.

Clara Bennett: That's where I land too, honestly. The rack-scale platforms are real. The benchmark is real. But custom ASICs growing nearly three times faster than GPUs — that's the signal that tells you where the volume is actually going. Nvidia defends a genuinely strong market. It's just a narrower one than the order pipeline implies right now.

Max Rivera: Yeah. I mean — three quarters from now we'll know a lot more about whether hyperscalers bought the kitchen or hired the architect. Either way, somebody's eating well. Thanks for working through this with me.

Clara Bennett: Good conversation. Genuinely.