Onpode
Cover art for Google bakes Gemini AI model directly into server chip for vastly more efficient inference

Google bakes Gemini AI model directly into server chip for vastly more efficient inference

July 20, 2026 · 9 min

Walt Garner & Nina Park

Google's internally codenamed 'Frozen v2' chip reportedly hardwires Gemini's neural-network architecture directly into silicon, projecting 6–10× more tokens per watt than current TPUs. The architecture — not the model weights — is locked. A 2028 target is cited, but Google has not confirmed the project exists.

Google is developing an experimental server chip, internally codenamed "Frozen v2," that embeds portions of its Gemini AI model's neural-network architecture directly into the chip's physical silicon circuitry. The project was first reported by The Information in July 2026 and subsequently covered by Reuters, CNBC, Bloomberg, and other outlets; Google has not officially confirmed it.

0:008:34
Make your own on Onpode

Describe any topic. Hear it in minutes.

More Onpode episodes on Technology

About this episode

Google is reportedly building a chip called Frozen v2 that does something no major hyperscaler has tried at this scale: it etches one specific AI model's structural architecture — Gemini's — directly into the silicon. Engineers cited in The Information's reporting project 6 to 10 times more tokens per watt than Google's latest TPUs. Google has confirmed none of it. This episode works through what that actually means, and why the distinction between freezing a model versus freezing an architecture matters enormously for how you evaluate the risk. The weights can still be updated. What cannot change is the structural blueprint — how the chip processes a token at the hardware level. That's what's locked. The episode also pushes back on the 'moonshot' framing that's dominated coverage. Amazon's Inferentia2 was already hitting comparable efficiency gains in 2023. Microsoft's Maia is in the same space. Google's 3.7% stock jump on an unconfirmed report is perhaps the most honest signal in the story — investors weren't rewarding ambition, they were exhaling because Google Cloud had been turning away paying customers and the compute crunch was visible from outside the building. The harder question the episode sits with: is there any public evidence that Google's own research roadmap for Gemini will stay architecturally stable through 2028? Because if it doesn't, the purpose-built silicon becomes a very expensive monument to a model that moved on.

Frequently asked

What is Google's Frozen v2 chip?

Google's Frozen v2 is a reported inference chip that hardwires Gemini's neural-network architecture directly into silicon circuitry. Unlike general-purpose TPUs, its structural blueprint is fixed at manufacture. Engineers project 6–10× more tokens per watt than Google's latest TPUs. Google has not publicly confirmed the chip's existence as of mid-2026.

Does baking Gemini into a chip mean the AI model can never be updated?

No — Google's Frozen v2 chip locks Gemini's architecture, not its weights. The learned parameters can still be updated, meaning the model can be retrained or fine-tuned. What cannot change is the structural blueprint: the way Gemini processes a token is permanently etched into the chip's transistors.

How does Frozen v2 achieve better inference efficiency than TPUs?

Frozen v2 eliminates the memory-shuttle overhead that TPUs incur on every inference run — the constant data movement between memory and compute. It also cuts redundant per-token calculations because the chip's circuitry already matches Gemini's exact computational shape, removing the performance penalty that general-purpose hardware pays for flexibility.

Why did Alphabet's stock rise on the Frozen v2 report?

Alphabet's stock rose approximately 3.7% in July 2026 following The Information's report on Frozen v2 — despite no confirmation from Google. Analysts attributed the move to investor relief: Google Cloud had been turning away paying customers due to compute constraints, and the chip report suggested a credible supply-side response.

What is the biggest risk of hardcoding an AI model's architecture into a chip?

The core risk is architectural obsolescence. If Gemini's structural blueprint changes before Frozen v2's projected 2028 deployment — due to new attention mechanisms, sparse architectures, or competing model designs — the purpose-built silicon cannot be redirected. Unlike TPUs, Frozen v2 cannot be retasked to a Gemini successor with a different underlying architecture.

Grounded in 12 sources
2 things capping Monday's market — plus, Alphabet's new AI chip roadmap - CNBC · cnbc.com
Alphabet stock pops on report it's developing a more efficient AI chip - CNBC · cnbc.com
Google plans new chip to run Gemini models more ... · reuters.com
Google developing Gemini-specific chip called Frozen v2 · tech.yahoo.com
Google is working on a new AI chip designed to make Gemini more efficient - TechCrunch · techcrunch.com
Satya Nadella has issued a shocking warning to companies using AI - TechCrunch · techcrunch.com
Google Frozen chip: Gemini baked into the silicon · thenextweb.com
Apple to skip high-end M6 Mac chips in favor of AI-focused M7 line | Hacker News · news.ycombinator.com
Alphabet Stock Gains On Report Of Google's New 'Frozen ... · tradingview.com
Hyperscaler XPUs: TPU, Trainium/Inferentia, Maia, MTIA · The Definitive Guide to AI Data Centers · aidatacenterguide.com
Ironwood TPUs and new Axion-based VMs for your AI workloads | Google Cloud Blog · cloud.google.com
Google designs new AI chip that bakes Gemini directly into silicon · cryptobriefing.com
Read transcript

Walt Garner: Nina, good morning — I assume your feed looked the same as mine this week.

Nina Park: Oh, complete chaos — and I'm glad, because I need someone to help me figure out if I should be impressed or worried. Probably both.

Walt Garner: Both seems correct. So, what we have here is a chip — internal codename Frozen v2 — first reported by The Information in July 2026, since picked up by Reuters, CNBC, Bloomberg. The premise: Google embeds Gemini's neural-network architecture directly into the chip's circuitry. The architecture is frozen into hardware.

Nina Park: And Google has said — nothing. Zero. No confirmation this project even exists.

Walt Garner: Indeed, and yet Alphabet's stock climbed 3.7% on Monday. The entire move — billions in market cap — on unconfirmed reporting from a single outlet. That is the thing that should make everyone stop.

Nina Park: Because it means the market is not actually reacting to a chip — it's reacting to the possibility that Google finally has a credible answer to something. And that something is that Google Cloud has been turning away customers. Paying customers. Because they can't serve the compute.

Walt Garner: Which is the confession buried inside the headline.

Nina Park: So what we're trying to work through today is whether locking Gemini's architecture into silicon by 2028 is an act of confidence or — I mean, an act of genuine desperation dressed up as strategy.

Walt Garner: Now, before we call it desperation — the coverage is muddying something important, and I want to slow down on it. The headline keeps saying the model is frozen. It isn't. The architecture is frozen. Those are not the same thing. The weights — the actual learned parameters — can still be updated. What you cannot change is the blueprint. The structural logic of how Gemini processes a token is what gets hardwired into the transistors.

Nina Park: Wait, so it's not like — the chip isn't a photograph of one specific Gemini.

Walt Garner: Exactly right. Think of it this way. Imagine a kitchen — every counter, every burner, every drawer — designed around one chef's specific menu. The kitchen runs faster, wastes nothing, because nothing in it is wrong for the task. But the chef can still tweak a sauce, adjust a seasoning. What they cannot do is decide to pivot to sushi. The counters are in the wrong place. The knives are wrong. Frozen v2 is the kitchen. Gemini's architecture is the menu baked into the countertops.

Nina Park: Okay that actually landed. So the efficiency — the six to ten times more tokens per watt than Ironwood — that number comes from the kitchen being exactly right, not from some magic.

Walt Garner: Yes, and Bloomberg specifically reported on the mechanism. You eliminate the memory-shuttle overhead — the constant back-and-forth of moving data around — and you cut redundant calculations per token, because the chip already knows, structurally, what shape the computation takes. You are not paying the penalty of generality. TPUs pay that penalty every single inference run.

Nina Park: And TPUs keep running everything else — training, other models, the stuff that needs flexibility.

Walt Garner: That is the official framing, yes — Frozen v2 as a specialized branch, not a replacement. Which is where I'd push back on the coverage, actually. The efficiency gain is real. The cost is that you have now built infrastructure that is architecturally locked to Gemini as it exists today, and deployment is 2028. That is — I mean, the transformer architecture felt revolutionary maybe five years ago, and the landscape has already shifted significantly since.

Nina Park: So the bet isn't just 'will this chip work.' The bet is 'will Gemini's blueprint still be the right blueprint when the kitchen opens for service.'

Walt Garner: And Google has not confirmed a single word of this. Everything — the 6–10x figure, the 2028 target, the architecture lock — all of it is The Information's reporting. That is worth holding.

Nina Park: Yeah, but holding that matters less to me right now than the framing problem — because the coverage keeps calling this a bold engineering moonshot. And I don't buy that. Google Cloud was turning away paying customers. That's not a moonshot. That's a supply crisis with a press release.

Walt Garner: Well, in fairness — those two things aren't mutually exclusive. A supply crisis can force genuinely good engineering decisions.

Nina Park: Okay but — actually, no, here's where I want to press on that — Amazon's Inferentia2 shipped in 2023. Already hitting eight to nine times better inference efficiency per watt over general-purpose GPUs. Google is not leading this. They're the last hyperscaler to admit they have the problem.

Walt Garner: That is — yes, and Microsoft's Maia is in the same category. The industry shift away from Nvidia's general-purpose GPUs is already well underway. Google arriving now is not a visionary pivot. It is a catch-up move.

Nina Park: Which is why the 3.7% stock jump is actually the most honest data point in this story — the market wasn't reacting to genius. It was reacting to relief. Like, finally.

Walt Garner: On an unconfirmed report.

Nina Park: On an unconfirmed report! That's the tell. Alphabet moves three-point-seven percent because investors have been watching Google Cloud turn away revenue, and the desperation is legible from outside the building.

Walt Garner: So the reframe you're making is — Frozen v2 isn't a bet on the future, it's a confession about the present. The compute crunch wasn't a background condition. It was the forcing function.

Nina Park: Exactly — and the part that makes this genuinely uncomfortable comes later, because we haven't even touched what happens to all that purpose-built silicon if Gemini's architecture looks different by the time 2028 actually arrives.

Walt Garner: And that is, I think, the actual wager nobody's naming directly. You see, there's a historical arc here worth tracing — CPUs gave way to GPUs because graphics became the dominant workload, GPUs gave way to TPUs because neural networks became the dominant workload. Each transition traded flexibility for performance on a narrower task. But Frozen v2 pushes that specificity further than any prior generation — it's not locked to AI workloads, it's locked to one model family's architecture. Gemini's. Specifically.

Nina Park: Which is a very different kind of bet than Trillium or Ironwood ever were.

Walt Garner: Considerably different, yes. And the thing Google has not explained — not once, publicly — is what internal evidence makes the stability assumption rational rather than wishful. The 2028 target is stated. The architecture lock is stated. The justification for believing Gemini's blueprint stays stable across that window is — well, it's absent.

Nina Park: Picture this — it's March 2028, Frozen v2 actually ships on schedule. But six months earlier, Google's own research team published a paper on a new attention mechanism that makes Gemini's current architecture look inefficient. Now you have infrastructure teams sitting on purpose-built silicon optimized for yesterday's model. What do they do? Because sunk-cost logic says keep running Frozen v2. But the better model is right there.

Walt Garner: That is — and I mean this precisely — that is not a hypothetical. Mixture-of-experts, sparse architectures, the attention mechanism experiments already published in the last eighteen months. The model landscape has shifted meaningfully within windows far shorter than two years.

Nina Park: And Google's own TPU family went through Trillium and then Ironwood before this question even arose. That's two generational pivots inside the same organization.

Walt Garner: Right — but the part that doesn't fit is that those were general-purpose evolutions. You could redirect a Trillium. You cannot redirect Frozen v2 to a Gemini successor with a different structural blueprint. That's the asymmetry.

Nina Park: So what do we actually watch for. Concretely.

Walt Garner: Two things. Any public signal from Google about Gemini's architectural roadmap — if they start describing the architecture as stable or canonical, that's meaningful. And whether 2028 slips. A delay isn't just an engineering problem, it's an admission that the model the chip was built for has already moved.

Nina Park: And that's the thing I keep sitting with — Google hasn't confirmed any of this. Not the chip, not the 2028 date, not the six-to-ten times figure. The Information reported it. The market rewarded it. And we're all here doing the math on a bet that Google hasn't even admitted it's making.

Walt Garner: Which brings you back, I think, to the question you framed at the start — and I'm genuinely not sure it's answerable from outside the building. Is the silence confidence, or is it the silence of a company that doesn't yet know whether what it's building will survive contact with its own research team.

Nina Park: Yeah. I don't think Google knows the difference right now either. And maybe that's — I mean, maybe that's the most honest thing about this whole story.

Walt Garner: Mm. Confidence and desperation, converging on the same decision.

Nina Park: Good talk. I'm going to be thinking about that kitchen metaphor for longer than I'd like to admit.

Google bakes Gemini AI model directly into server chip for vastly more efficient inference · Onpode