Onpode
Cover art for Why safety and capability research pull in opposite directions with different reward structures

Why safety and capability research pull in opposite directions with different reward structures

August 9, 2026 · 11 min

David Sterling & Megan Skiendel

AI capability research attracts far more funding than safety research because capability can be scored on public leaderboards — directly tied to enterprise procurement and government contracts — while alignment progress cannot. No equivalent safety leaderboard exists, creating a structural incentive gap that voluntary frameworks like NIST's AI Risk Management Framework cannot close.

The AI field exhibits a structural incentive asymmetry between capability research and alignment research that is rooted in economic and competitive dynamics rather than temporary policy gaps.

0:0011:18
Get the next episode on Artificial Intelligence

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Artificial Intelligence

About this episode

The gap between AI safety research and capability research isn't just a funding imbalance — it's a measurement problem with a structural incentive baked in. This episode traces how that structure formed and why it's so hard to dismantle. It starts with ImageNet in 2010: a public benchmark that made computer vision R&D legible to investors almost overnight. The same logic runs through AlphaGo, AlphaStar, and today's frontier leaderboards, where positions are directly tied to enterprise deals and government contracts. Capability gets measured, so capability gets funded. Alignment accumulates as an externality nobody's pricing. The episode digs into why even well-intentioned labs end up inside this structure. Alignment findings get published openly — they're a public good by design, which means they produce no competitive moat. The resource split between capability training and safety work isn't disclosed in public technical reports. And voluntary frameworks from NIST and the OECD, however serious, don't create a positive reward — they're restrictions without a price signal attached. The most useful frame the episode offers: the historical parallel to environmental regulation. What actually moved corporate behavior wasn't voluntary guidelines. It was liability moments and procurement requirements that internalized costs firms had previously exported to society. The question isn't whether labs are bad actors. It's whether the external pressure arrives before something makes the cost undeniable in the worst way.

Frequently asked

Why does AI capability research get more funding than AI safety research?

AI capability research gets more funding because benchmark scores on public leaderboards directly influence enterprise procurement deals and government contracts. Alignment progress, by contrast, cannot be reduced to a comparable public number, so it remains a narrative that investors and boards cannot price — making it structurally underfunded relative to capability work.

What is the 'alignment tax' in AI development?

The alignment tax refers to the competitive disadvantage safety-focused AI labs face because capability is the only thing the market has learned to price. Labs like Anthropic that prioritize safety still compete on public leaderboards, because there is no procurement signal that rewards alignment investment — making safety a cost center rather than a source of competitive advantage.

Can AI safety and capability research advance at the same time?

Researchers such as Eric Chu at MIT Media Lab argue that 'intersection research' — projects advancing both safety and capability simultaneously — is possible and undermines the assumption that the two are inherently in conflict. However, Chu himself acknowledges this approach remains a minority research philosophy at frontier labs and has not shifted field-wide resource allocation.

Would government procurement requirements actually improve AI safety?

Government procurement requirements are the mechanism most likely to shift AI safety investment, according to analysis of the current incentive structure. If agencies such as GSA or DoD conditioned contracts on a scored safety standard — analogous to FedRAMP for cloud security — safety metrics would become a market entry requirement rather than a voluntary cost center, creating a genuine price signal.

Why don't voluntary AI safety frameworks like NIST's AI Risk Management Framework fix the problem?

Voluntary frameworks like NIST's AI Risk Management Framework set behavioral guardrails but create no positive financial reward for safety investment. No firm internalizes a cost it is not legally required to carry. Historical analogues — environmental regulation, asbestos liability — show that field-wide behavior changes only when the external cost of harm is made internal through liability rules or mandatory procurement standards.

Grounded in 9 sources
Capabilities ∩ alignment research — Eric Chu · alumni.media.mit.edu
How should AI Safety Benchmarks Benchmark Safety? · arxiv.org
Mind the Gap! Pathways Towards Unifying AI Safety and Ethics Research · arxiv.org
From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution · arxiv.org
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding · arxiv.org
From Principle to Practice: Value Alignment in AI Ethics and Governance | German Law Journal | Cambridge Core · cambridge.org
From benchmarks to deployment: a comprehensive review of agentic AI ... · link.springer.com
Industrial policy for the final frontier: Governing growth in the emerging space economy | Brookings · brookings.edu
Tipping the Cyber Balance: How AI Benchmarks Could Make Software Safer | RAND · rand.org
Read transcript

Megan Skiendel: Hey — before we get into anything, tell me if this image works for you. Stadium. Scoreboard goes up for the first time. Every player on the field immediately starts optimizing for whatever the scoreboard measures, whether or not that's actually the game.

David Sterling: I mean — that's ImageNet in 2010. Exactly.

Megan Skiendel: Which is where today starts. Because that scoreboard moment, that 2010 ImageNet launch, is what made a decade of computer vision R&D fundable. Not better vision algorithms — a public number labs could race toward.

David Sterling: And now every frontier lab — OpenAI, Google DeepMind, Anthropic, DeepSeek — they are all playing the same game on public AI leaderboards. Those positions are directly tied to procurement money. Enterprise deals, government contracts.

Megan Skiendel: Millions of dollars tracking a leaderboard score.

David Sterling: Right — but here's the point. There is no equivalent leaderboard for AI alignment. Not one that moves procurement decisions. Not one that Google DeepMind's board is watching.

Megan Skiendel: And that absence — honestly, I don't think it's an accident. The measurement asymmetry is the mechanism. You can show a benchmark score to an investor in thirty seconds. Alignment progress is a narrative that takes months to build, and even then you can't put it on a leaderboard.

David Sterling: Which is the structural incentive misalignment in one sentence: capability gets measured, therefore capability gets funded, and safety accumulates as an externality nobody's pricing.

Megan Skiendel: So the question driving everything today — is this a policy mistake someone could fix? Or is it just how markets work when you can't measure the thing that matters?

David Sterling: And that's exactly where AlphaGo becomes the template, not just an example. When Google DeepMind beat Lee Sedol in March 2016 — four games to one — that wasn't a press release. That was a pricing signal. Every government AI office, every enterprise R&D budget, every VC portfolio shifted toward whatever AlphaGo represented.

Megan Skiendel: A single match reshapes the funding landscape.

David Sterling: Four to one. And then AlphaStar — 2019, dominating professional StarCraft II players — same mechanism, second iteration. The field learns: win publicly, win visibly, and the capital follows. That template is now just how the industry works.

Megan Skiendel: But both of those are capability milestones. Nobody put up a scoreboard for — I don't know — 'how well does this system refuse a harmful request.'

David Sterling: Exactly. Because refusing a harmful request doesn't beat a world champion on camera. Now look at the GPT-4 Technical Report — OpenAI published it in 2023, and it's actually useful as evidence here. They documented RLHF, safety testing, alignment work — all in the same document as the capability training. Which sounds like balance.

Megan Skiendel: But if they're doing both, isn't that — I mean, isn't that good enough?

David Sterling: That's the wrong question. Doing both is not the same as doing both at equal scale. And the report doesn't publish the budget split. The compute allocated to capability training versus the compute for RLHF. We don't have that number. The asymmetry is structural and it's also deliberately opaque.

Megan Skiendel: Wait — the report exists partly to signal that they're responsible, but it doesn't actually let you verify the resource split?

David Sterling: Correct. RLHF runs alongside large-scale capability training — not instead of it. That's OpenAI's own framing. Alignment supplements the capability pipeline. It doesn't constrain it. And we can't tell how wide the gap is because the balance sheet isn't public, which is — frankly — itself a data point.

Megan Skiendel: The absence is information.

David Sterling: The absence is the signal. And this is why benchmark position functions as a market mechanism, not just a research scorecard. When DeepSeek posts a leaderboard result that competes with GPT-4 level performance — that immediately moves enterprise procurement conversations. No equivalent exists for alignment. There's no procurement officer saying 'I'll pay a premium because your RLHF budget was forty percent of compute.' That market doesn't exist yet.

Megan Skiendel: Which means the alignment tax is real — not because safety makes models worse, but because capability is the only thing the market has learned to price. And Anthropic is — honestly, they're living proof of that problem. They built the lab explicitly around safety being the core investment, and they're still fighting for market share against labs that just — move faster on the leaderboard.

David Sterling: But that Anthropic point actually cuts both ways — and it exposes the deeper problem. It's not just that the market doesn't reward safety. It's that you literally cannot show alignment progress on a slide.

Megan Skiendel: Say more. What does that actually look like inside a lab?

David Sterling: Picture a lab VP. Thursday afternoon, board meeting the next morning. She has two slides. Slide one: 'Model jumped four points on MMLU — here's the graph, clean upward line.' Slide two: 'Alignment team believes the model is somewhat less likely to produce harmful outputs in edge cases — here is a qualitative report.' Which slide gets the follow-on funding commitment?

Megan Skiendel: Slide one. Every time.

David Sterling: Every single time. Not because the board is evil. Because one is a number and one is a narrative. That's the measurement asymmetry doing actual work — it's not a background condition, it's the decision mechanism.

Megan Skiendel: And the academic publishing side makes it worse. Novel capability results — demonstrable, impressive, reviewers can evaluate them in an afternoon. Alignment work is incremental, hard to evaluate, and honestly? Harder to get funded in the first place because the outputs look weak on paper.

David Sterling: Which compounds the competitive moat problem. A capability breakthrough — you can retain it, differentiate your product, build first-mover advantage. Alignment findings? Labs publish them openly. Anthropic publishes, OpenAI publishes. No proprietary advantage. So the economic logic of even doing the work is — I mean, it collapses.

Megan Skiendel: Wait — so the more responsible you are, the less competitive value you extract from the safety work you actually did?

David Sterling: That's the structure, yes. Published alignment research is a public good. Publicly shared. Which is — frankly — the right thing to do. But it produces zero moat.

Megan Skiendel: There is a counterargument here though — and I think it's worth taking seriously. Eric Chu at MIT Media Lab has been mapping what he calls intersection research. Capabilities intersect with alignment — projects that advance both simultaneously. His argument is the two don't have to be orthogonal. That the alignment tax framing is actually a choice labs make, not a physics law.

David Sterling: I'd want to see the proportion. If intersection research is — what — ten percent of active projects at a frontier lab? That's a minority philosophy, not a structural fix.

Megan Skiendel: Chu himself says it remains a minority research philosophy. He's not claiming it's dominant. He's saying it's possible — and that matters because right now the labs treat capability and alignment as if they're in direct conflict by default, and that assumption does actual organizational damage.

David Sterling: The assumption shapes the headcount. Which shapes the budget. Which shapes what gets measured — and we're back to the same loop.

Megan Skiendel: And look — the part that I think is actually going to break people's brains is where this goes next. Because NIST and the OECD have both proposed structural fixes. Anthropic has a public position on this. Whether any of that moves the needle on a problem that's fundamentally about what gets measured — that's a harder case to make than it looks.

David Sterling: And that's exactly where the NIST and OECD frameworks run into the wall. Because those proposals — they're behavioral guardrails. Thou shalt not deploy without a risk assessment. But the structural problem isn't bad behavior. It's the absence of a positive reward for safety. Nobody gets paid more for passing a NIST checklist.

Megan Skiendel: So restriction without reward.

David Sterling: That's the gap. And the historical analogy actually makes this precise. Environmental damage, public health harms — firms underinvested for decades. Not because they were malicious. Because the cost fell on society, not the firm. Same structure. And what actually moved it? Not voluntary frameworks. Liability rules. Procurement requirements. Mechanisms that made the external cost internal.

Megan Skiendel: And we don't have the AI equivalent of that liability regime yet.

David Sterling: Not even close. OECD guidelines are voluntary. NIST's AI Risk Management Framework — useful, serious work — but it's a framework, not a liability rule. No firm internalizes a cost they aren't legally required to carry.

Megan Skiendel: Which is — honestly, that's where Anthropic becomes the strangest case study. They were founded to make safety the core line item. That was the explicit founding logic. And they compete on public leaderboards.

David Sterling: That surprises you?

Megan Skiendel: It should surprise everyone. Because if an organization built from scratch specifically to resist this structure still ends up inside it — then identity alone is not a corrective mechanism.

David Sterling: That's the load-bearing point. Individual lab leaders — yes, they can make unilateral safety investments that defy pure economic logic. Some have. But field-wide resource allocation hasn't moved. One org's culture doesn't reprice the market.

Megan Skiendel: So what would actually shift it? Because — I mean, the regulation answer feels incomplete after what you just laid out.

David Sterling: The honest answer — and I'll give you the honest one — is we don't know. What the environmental analogy predicts is you need the liability moment first. A specific failure, attributable, costly enough to change the cost structure firms face. Before asbestos litigation, manufacturers did not voluntarily exit the market. After it, the calculus changed overnight.

Megan Skiendel: We're waiting for the asbestos moment in AI.

David Sterling: Or a procurement shift — government says 'we will only buy models that meet a specific, scored safety standard.' That creates a positive reward. That's the mechanism. Not a checklist. A price signal that flows through to the board slide on Thursday night.

Megan Skiendel: And right now that slide doesn't exist. Which means the structural externality just — compounds. Every quarter, capability benchmarks get more granular, more public, more tied to procurement. And the alignment side of the ledger stays a narrative nobody can price.

David Sterling: The procurement angle is the one mechanism I'd actually watch. If the federal government — GSA, DoD, whoever writes the contract — conditions purchasing on a scored safety standard, that's not a restriction. That's a price signal. Suddenly safety metrics are an input to winning, not a cost center.

Megan Skiendel: And that flips the board slide. Thursday night, the graph that matters is suddenly — what, 'our safety score qualifies us for the contract pool.'

David Sterling: That's the structural flip, yes. Alignment becomes a path to competitive moat — not a public good you donate to the field. Which is the only version of this that doesn't require labs to act against their economic interest.

Megan Skiendel: Honestly, and I've sat in rooms where people have floated exactly that — frame safety certification the way you'd frame FedRAMP compliance. You don't get the government cloud contract without it. Nobody complains that FedRAMP is a cost center. It's a market entry requirement. But — I mean, the gap is that FedRAMP measures something auditable. We don't have the safety equivalent of an audit standard yet. NIST's framework is serious, OECD has proposed guidelines, but neither one is a scored threshold you either clear or you don't.

David Sterling: Which means the procurement mechanism and the measurement problem are actually the same problem. You can't condition a contract on a score that doesn't exist.

Megan Skiendel: And that's — I think that's actually where I land on all of this. The thing I keep returning to isn't whether the labs are bad actors. It's that the market genuinely does not have a price for safety progress. No leaderboard, no procurement signal, no moat. And until that changes — from the outside, through policy, or through a failure that makes the cost suddenly visible — the structure just holds.

David Sterling: The open question being whether that external pressure arrives before something makes the cost undeniable in the worst way.

Megan Skiendel: Yeah. Not a comfortable place to stop, but an honest one. Thanks for walking through the structure with me — this one needed the numbers behind it.