Onpode
Cover art for The structural misalignment — capabilities scale predictably, alignment doesn't

The structural misalignment — capabilities scale predictably, alignment doesn't

September 22, 2026 · 14 min

Iris Holm & Lila Soto

AI capability scales predictably via power-law formulas — Kaplan et al. (2020) and Chinchilla (2022) let labs forecast model performance from compute and data budgets. AI alignment has no equivalent formula: it remains judgment-intensive, human-dependent, and non-compounding. That structural gap — not a lag — widens with every capability jump.

AI capability improvements follow predictable scaling laws: as researchers increase model parameters, training data volume, and computational power, language model performance declines in error according to smooth power-law relationships. OpenAI's Kaplan et al. (2020) established these relationships, showing that performance could be forecast from smaller experiments.

0:0013:46
Get the next episode on Artificial Intelligence

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Artificial Intelligence

About this episode

There's a formula for making AI more powerful. Kaplan et al. published it in 2020: language model error follows a smooth power-law function of parameters, training data, and compute. You can run a small experiment and forecast how a model ten times bigger will perform. Hoffmann et al. refined the ratios in 2022 — most frontier labs were over-indexing on parameter count and starving models of data — but the framework held. The argument was quantitative. Someone could run the numbers. This episode asks a harder question: where's the alignment equivalent? The formula for making systems behave according to human intentions as they get more capable? The honest answer is that there isn't one. Red-teaming, RLHF, interpretability research — all genuine advances, all labor-intensive, none of them compounding the way adding parameters compounds. The Alignment Forum is where much of this work gets hashed out: distributed, ad hoc, the structural opposite of an industrial pipeline. What the episode works through carefully is why this is a structural gap, not a lag. More capable systems produce novel failure modes that prior benchmarks never anticipated. Every release cycle resets the scope of alignment work. And the market structure makes it hard to fix: capability breakthroughs generate funding, alignment is a cost center. The decision about pace defaults to whoever is scaling fastest. No institution — not NIST, not any lab — currently has the authority to change that. A clear-eyed look at what it would actually take.

Frequently asked

What are AI scaling laws and why do they matter?

AI scaling laws, established by Kaplan et al. at OpenAI in 2020, show that language model error drops as a predictable power-law function of three variables: parameter count, training data volume, and compute budget. Labs can run small experiments, measure the slope, and reliably forecast how a model ten times larger will perform.

What is the difference between the Kaplan and Chinchilla scaling laws?

Both Kaplan (2020) and Chinchilla (Hoffmann et al., DeepMind, 2022) operate inside the same forecastable power-law framework but disagree on optimal ratios. Chinchilla showed that most frontier models were undertrained — over-indexed on parameter count while starving on data. For a fixed compute budget, balancing model size and token count outperforms maximizing parameters alone.

Why doesn't AI alignment scale the way AI capabilities do?

AI alignment — making systems reliably follow human intentions — has no equivalent scaling law. Techniques like RLHF, red-teaming, and interpretability research are all human-dependent and judgment-intensive. None compounds automatically when a model grows larger. More capable models also introduce novel failure modes that prior benchmarks never anticipated, continuously resetting alignment's scope.

What is AI red-teaming and why can't it keep up with capability scaling?

AI red-teaming involves human security researchers probing models for harmful behaviors — NIST's AI Risk Management Framework specifically names scenarios like enhanced phishing as required test cases. But red-teaming relies on human instinct and domain expertise; it cannot be automated or accelerated by adding compute. It does not compound the way capability research does.

Is the AI capability-alignment gap just a temporary lag that will close over time?

The capability-alignment gap is structural, not a temporary lag. Capability research is algorithmic, forecastable, and self-funding — benchmark gains justify further investment. Alignment is open-ended, context-dependent, and treated as a cost center. Larger models generate novel failure modes faster than alignment methods can track them, so the gap does not naturally close as time passes.

Grounded in 12 sources
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws · arxiv.org
Institutional AI: A Governance Framework for Distributional AGI Safety · arxiv.org
AI Alignment: A Comprehensive Survey - arXiv · arxiv.org
Automated Alignment is Harder Than You Think · arxiv.org
Deliberative Technology for Alignment · arxiv.org
Reconciling Kaplan and Chinchilla Scaling Laws · arxiv.org
From Principle to Practice: Value Alignment in AI Ethics and Governance | German Law Journal | Cambridge Core · cambridge.org
Understanding Kaplan and Chinchilla Scaling Laws in Simple Terms · medium.com
Teaching Claude Why - Alignment Science Blog · alignment.anthropic.com
Full automation of AI R&D probably yields a large speed up even without a software-only singularity — AI Alignment Forum · alignmentforum.org
The current bottleneck is political will, not research — AI Alignment Forum · alignmentforum.org
AI for AI safety — AI Alignment Forum · alignmentforum.org
Read transcript

Lila Soto: Can I just name something that's been bothering me before we do anything else — I promise it's relevant.

Iris Holm: Sure, that's usually where the good stuff is.

Lila Soto: We talk about AI risk like it's a vibe. Like, generally concerning, direction unclear. But then I looked at what Kaplan et al. published at OpenAI in 2020 — a paper called Scaling Laws for Neural Language Models — and it's the opposite of a vibe. It's a formula. Language model error follows a power-law function of parameters, training data, compute. That's industrial-grade predictability.

Iris Holm: Right. So where's the bothering part?

Lila Soto: The bothering part is — I went looking for the alignment equivalent. The formula for making it safe as it gets more powerful. And I mean... there isn't one. Not even a rougher version.

Iris Holm: That's not a coincidence. That's the whole problem. Capability follows scaling laws. Alignment — getting systems to actually behave according to human intentions — doesn't. It's context-dependent, it requires judgment, it's not automatable.

Lila Soto: Huh — and is that a known known, like inside the labs? Or is it more sort of... uncomfortable background noise?

Iris Holm: NIST made it a known known when they published their AI Risk Management Framework and specifically named malicious code generation and enhanced phishing as the scenarios you have to test against. The answer they gave: structured, human-driven adversarial testing. Red-teaming. Which is exactly the kind of thing that doesn't scale the way adding a billion parameters scales.

Lila Soto: So one side of this — making it stronger — is basically industrialized. And the other side is still people in rooms trying to break it by hand.

Iris Holm: Now you've got it. That asymmetry is structural. It doesn't resolve just because capability keeps compounding.

Lila Soto: That asymmetry is what I want to pull on, actually. Because saying 'capability scales, alignment doesn't' — I get that structurally. But I don't know what it *feels* like to be inside the machine that's doing the scaling. Like, what is a scaling law, in practice, to someone running the experiment?

Iris Holm: Think of it like a recipe. Except someone worked out the exact ratios — and it holds. More of this ingredient, performance goes up by this much. Predictably. Every time.

Lila Soto: A recipe where you actually know the ratios.

Iris Holm: That's what Kaplan et al. published in 2020. Language model error — how wrong it is — drops as a smooth power-law function of three things: parameters, training data volume, compute budget. You can run a small experiment, measure the slope, and forecast how a model ten times bigger will perform. That's engineering. That's not guessing.

Lila Soto: But didn't someone come back two years later and say the recipe was actually wrong?

Iris Holm: Wrong ratio, not wrong framework. Hoffmann et al. at DeepMind — the Chinchilla paper, 2022 — showed that most frontier models were undertrained. Labs were, I mean, they were throwing compute at parameter count while starving the model of actual data. Parameter-heavy, data-light. The Kaplan prescription had overweighted size.

Lila Soto: Oh — so the whole GPT-3 era was optimizing the wrong variable?

Iris Holm: For a given compute budget, yes. Chinchilla-optimal training says: balance model size and training token count. Don't just maximize parameters. The disagreement between Kaplan and Hoffmann is real — but notice what kind of disagreement it is. It's quantitative. It happened *inside* the same forecastable framework. You can run the numbers both ways.

Lila Soto: The argument was about the recipe, not whether a recipe exists.

Iris Holm: Exactly — and it gets more granular than that. There's a further refinement: inference-cost-adjusted scaling. If you're going to deploy a model to millions of users, the *optimal* size actually shifts downward because inference is expensive. Deployment context changes the math. And the framework handles it. That's what I mean when I say this is engineering now.

Lila Soto: Wait — so someone at OpenAI or DeepMind can sit down before they train anything and calculate roughly what they're going to get?

Iris Holm: Within the capability dimension — yes. That's the thing. The capability curve keeps getting redrawn, but it's still a curve someone can draw. Alignment has no equivalent curve. Not a rougher one, not a slower one. That's not a lag. That's a different category of problem.

Lila Soto: But that's what I keep snagging on — no curve at all, not even a bad one. Like, picture someone on a red team. A security researcher, it's a Saturday afternoon, she's got a model open and she's just... trying things. Trying to get it to help her write a phishing email. There's no algorithm telling her where to probe next. It's just her instincts, her domain knowledge, her sense of what might slip through.

Iris Holm: That's red-teaming. And NIST named exactly that scenario — enhanced phishing — as a required test case in the AI Risk Management Framework. So the threat is institutionally named. But naming it doesn't scale the solution.

Lila Soto: Right — and she can't just run that test more efficiently by... adding more of herself.

Iris Holm: No algorithmic substitute for her judgment there. Full stop.

Lila Soto: Okay but — I mean, isn't RLHF supposed to be the answer to that? Like, Anthropic is working on this, constitutional AI is a real thing. The labs aren't just ignoring alignment.

Iris Holm: They're not ignoring it. RLHF — preference training, human evaluators rating outputs — it's a genuine advance. Constitutional AI is real. But both of them are human-dependent at the core. You still need people making judgments about which outputs are good. That doesn't compound automatically when the model gets bigger.

Lila Soto: So it's... labor-intensive the whole way down.

Iris Holm: And interpretability research is worse, structurally. You're trying to understand what's actually happening inside a neural network — the internal representations. That's a novel research domain. Insight doesn't accumulate just because you train a larger model. The Alignment Forum is where a lot of this gets worked out — distributed, ad hoc, researchers posting findings. It's the opposite of an industrial pipeline.

Lila Soto: The Alignment Forum — like, that's the infrastructure? Forum posts?

Iris Holm: That's not a knock on the researchers. It's a description of the structure. Capability research has a compounding funding loop — every benchmark gain justifies the next investment. Alignment has no equivalent. It's allocated as a fraction of capability budgets, not its own priority.

Lila Soto: So the structure of the problem resists the industrialization that capability achieved. That's... yeah, that's the uncomfortable version of it.

Iris Holm: And no alignment technique — not red-teaming, not RLHF, not interpretability research — has demonstrated the efficiency compounding that scaling laws provide on the capability side. None.

Lila Soto: Which makes me want to ask why that gap doesn't just... close over time. Like, is it a lag, or is it something structural that institutions genuinely can't fix — and I think that's the real issue people don't want to admit.

Iris Holm: Structural. Not a lag. And here's why that distinction matters — more capable systems don't just hit the old failure modes harder. They produce novel failure modes that prior benchmarks never anticipated. Ones that didn't exist to red-team against.

Lila Soto: So the denominator keeps moving.

Iris Holm: Faster than alignment can track it. Every capability jump resets scope for alignment work. That's not a temporary problem. The capability-alignment gap is a feature of what the two kinds of research actually are — capability research is algorithmic, forecastable, compounding. Alignment is open-ended, context-dependent, judgment-intensive. Those natures don't converge just because time passes.

Lila Soto: But don't institutions help close it? Like, NIST put out a whole framework. That's not nothing.

Iris Holm: NIST names the right activities. Red-teaming, adversarial testing — yes. But recommending red-teaming is not the same as red-teaming scaling at the rate capabilities do. The framework acknowledges the resource asymmetry without solving it. That's acknowledgment, not closure.

Lila Soto: Huh. So the governance layer is, I mean — it exists, but it's not operating on the same curve.

Iris Holm: Deliberative technology proposals are the clearest example. Democratic deliberation for alignment — you convene people, you hash out values, you try to specify what a system should and shouldn't do. That's entirely outside the scaling curve. It doesn't get faster when the model gets bigger.

Lila Soto: Okay but — there's a version of this story we reach for culturally. Nuclear energy, synthetic biology. Capability outpaced safety there too, and eventually mature safety regimes developed. Why isn't that just... the arc we're on?

Iris Holm: The dual-use analogy is genuinely useful up to a point. And then it imports false comfort. Nuclear and synthetic biology had physical constraints — you can't enrich uranium faster by wanting to, the biology has rate limits. And both fields had engineering precedents. Safety was hard, but the problem was bounded.

Lila Soto: And AI alignment isn't bounded the same way.

Iris Holm: Because — and this is the genuinely novel part — alignment requires specifying and verifying human values inside a learned system. Value specification, interpretability research, that's not an engineering problem with precedents. Nuclear safety never had to answer 'what does the reactor actually want.' Alignment does. Roughly.

Lila Soto: So we reach for that historical comfort and it kind of... almost fits. But the piece that doesn't fit is the hardest piece.

Iris Holm: Right. Picture a safety researcher at Anthropic — it's a Wednesday, late, she's got a new model version open. She's not testing old failure modes. She's probing for behaviors that nobody catalogued before this model existed. There's no checklist that covers it. That scope expands every release cycle. Without deliberate investment that matches that expansion, more capable systems become harder to align. Not easier. The trajectory runs the wrong direction.

Lila Soto: And the gap isn't about any single model's release date — it's baked into what these two types of research fundamentally are.

Iris Holm: Baked in. That's the right word. Kaplan et al. and Hoffmann et al. — both operating inside a quantitative, forecastable framework. You can disagree about the ratios, still run the numbers. Alignment has no numbers to run. That's not a funding problem you fix with a budget line.

Lila Soto: I think what shifted for me — and I'm still kind of sitting with it — is that I came in thinking this was a pacing problem. Like alignment is just slower, and eventually it catches up. But the two things have different shapes. Capability research compounds. Alignment accumulates, slowly, by hand, with human judgment. Those aren't the same shape at all.

Iris Holm: Different shapes. That's exactly it. And the implication that follows — quietly, without any drama — is that if alignment doesn't yield something like a scaling law in the next generation of frontier models, then someone has to decide whether to slow capability development. And right now there is no institution with that authority. Not NIST, not Anthropic, not OpenAI.

Lila Soto: No one's even really assigned to make that call.

Iris Holm: No one. And the market structure makes it actively hard to assign. Capability breakthroughs generate funding. Alignment doesn't generate funding the same way — it's a cost center, not a return. So the decision about pace defaults to whoever is scaling fastest.

Lila Soto: Which means the gap stays structural unless someone deliberately — and I mean deliberately, not as a side effect of anything else — invests in the slower kind of work. The red-teaming, the interpretability research, the RLHF labor. None of that speeds up on its own.

Iris Holm: None of it. And that's where I'll actually stop pushing, because I think that's the honest shape of it. Not alarming in the way that makes people tune out. Just — accurate.

Lila Soto: Yeah. The gap is real, it's structural, and the only thing that closes it is deliberate investment in the work that doesn't scale. That's kind of a modest conclusion for something this large.

Iris Holm: Modest conclusions are underrated. Thanks for dragging this one into the light.