Hugo Vance: You sent me that paper at eleven at night with no context — just 'read this.' I did.
Lila Soto: And? I mean — did it land?
Hugo Vance: It landed. The Chinchilla correction. Hoffmann et al. at DeepMind, using power-law curves — the same mathematical architecture Kaplan et al. used at OpenAI in 2020 — concluding that large models of that era were undertrained by something like a factor of twenty. That is a significant number.
Lila Soto: Twenty times. Not a rounding error — that's a systematic blindness baked into how the whole field allocated compute. And that's actually where I want to start today, because the question underneath all of this is: if scaling laws are laws, how does the same power-law math produce one prescription in 2020 and the opposite prescription two years later?
Hugo Vance: The laws contradicting themselves.
Lila Soto: Yeah — and Jared Kaplan, the OpenAI author who gave the field the first recipe, goes on to co-found Anthropic, which is now one of the labs most committed to scaling. So the person whose paper got revised is now building the thing the revision says should have been built differently all along. That's not just interesting, it's kind of the whole tension.
Hugo Vance: And yet before we can hold that tension — the same math, two contradictory prescriptions — I think we need to actually say what a scaling law is. Not the formula. The intuition. Think of it like a mileage curve on a very long road trip. You don't need to drive the whole highway to know your rate. The first fifty miles tell you. And that rate — it holds. That's what Kaplan et al. at OpenAI established in 2020: on a log-log plot, cross-entropy loss — how well the model predicts the next token — improves approximately linearly as you add parameters, data, compute. Smooth. Predictable. Many orders of magnitude.
Lila Soto: So a researcher reading that paper, late 2019, deciding whether to build one 175-billion-parameter model or spread the compute — the curve says go big on parameters?
Hugo Vance: The Kaplan recipe said exactly that. Favor parameters; data matters less at the margin. She builds the large model. And then Hoffmann et al. at DeepMind run the same power-law math — and find she starved that model of training tokens by a factor of roughly twenty. Compute-optimal allocation, they call it: parameters and training tokens should scale in roughly equal proportion. The curve was real. The reading of the curve was wrong.
Lila Soto: Twenty times undertrained. That's — I mean, that's not a calibration issue.
Hugo Vance: No. And here's where I'd be cautious about calling this a law correcting itself. A law doesn't produce opposite prescriptions from identical inputs. What Chinchilla scaling revealed is that the original framework optimized the wrong variable — it fit loss curves as a function of model size and missed the interaction with data volume entirely. That's not self-correction. That's a gap in the original model.
Lila Soto: Hm — so the practical payoff, capability forecasting, extrapolating these log-log curves before you commit a compute budget — that whole enterprise depends on having fit the right variables in the first place?
Hugo Vance: Yes. And there's a harder version of that problem coming — because the loss curves stay smooth even when the model's actual behavior does something the curve didn't predict. We'll get to that. The predictability is real, and it may also be hiding something.
Lila Soto: And that hiding — that's actually the thing that won't let me go. Because Daniela Amodei and Anthropic's whole public position is basically: we know roughly what we're buying when we scale. The curves have held. But then you get emergent capabilities — model does something it literally could not do before, at some threshold of scale, and the loss curve the whole time? Smooth descent. No bump. No warning.
Hugo Vance: That's the crack in the glass.
Lila Soto: Picture someone at Anthropic — it's, I don't know, a Friday afternoon, they're reviewing eval results from a new checkpoint — and cross-entropy loss is down, exactly on the curve, everything nominal. And then a benchmark they'd been tracking as flat just... lights up. Multi-step reasoning, something it couldn't do at all. The curve didn't predict that. The curve measured the wrong layer of what was changing.
Hugo Vance: Yes. And I think we have to be careful not to overstate — the sourcing on emergence is genuinely contested. But the gap between 'cross-entropy loss improved' and 'a qualitatively new capability appeared' is real and unresolved. Those are two different questions. The framework answers one cleanly and is silent on the other.
Lila Soto: So is it measuring the wrong thing when it matters most?
Hugo Vance: That is the live question. And it's why the formal 'Enough of Scaling LLMs' camp — the downscaling position — has traction. Because if emergence is real and unpredicted by the curves, then inference-time compute, efficiency, different axes entirely, those aren't retreats. They're bets that the three-variable model — parameters, data, compute — is already an incomplete map.
Lila Soto: Yeah — and I don't think either of us can say confidently which bet is right. The curves have been robust across architectures and domains, which is why universality claims keep getting made. But robust within a regime isn't the same as universal. That gap is still open.
Hugo Vance: Well. Stanford HAI and the Kempner Institute at Harvard are actively rewriting pieces of the framework — learning rate schedules, data distribution effects, configuration-level corrections. You can read that as reassuring. The science is honest, it's stress-testing itself. Or you can read it as the map keeps changing, which is — I mean, that is not nothing, when billion-dollar training runs are navigating by that map.
Lila Soto: Yeah — and I think that's what I'll actually carry out of this. 'Scaling law' is almost a misnomer. It's a working model the field keeps pressure-testing, and the pressure tests keep finding edges. The bet isn't that the curves hold forever. It's on whoever understands what the curves can't see.
Hugo Vance: You sent me that paper at eleven at night. No context. And I think what I didn't expect — it's not that the correction was large. It's that the correction was possible at all inside the same framework. That's either the health of the thing or its exposure. I'm genuinely not sure which.
Lila Soto: Good. I'll take genuinely not sure. Thank you for sitting with it.