Onpode
Cover art for In silico antibody design faces the reality check of binding assays

In silico antibody design faces the reality check of binding assays

August 24, 2026 · 8 min

Maya Chen & Dr. Nathan Hayes

The AIntibody benchmark — organized by Specifica LLC and published in Nature Biotechnology — tested 511 antibody sequences from 29 organizations in blinded, prospective wetlab assays. Inverse folding models using AlphaFold 3 structures achieved the highest Spearman correlation and best precision@10, yet structural accuracy still did not reliably predict binding affinity or developability.

The AIntibody benchmark is a blinded, prospective challenge designed to evaluate whether AI-generated antibody candidates can be discovered computationally and then confirmed through experimental measurements of affinity and developability.

0:008:27
Get the next episode on Science

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes

More Onpode episodes on Science

About this episode

For years, the implicit promise of AI antibody design was that if you could predict the structure well enough, everything downstream would follow. AIntibody — a blinded, prospective benchmark organized by Specifica LLC — just tested that promise against experimental reality, and the results are more specific, and more unsettling, than a simple pass or fail. This episode works through what it actually meant to run 29 organizations through a challenge where no one could revise their predictions after seeing the wetlab data. Inverse folding models using AlphaFold 3 structures led on Spearman correlation and precision@10 — a real, measurable advance. But the episode traces why that doesn't close the loop: structural accuracy and binding affinity are genuinely different problems, and the benchmark made that gap impossible to paper over. The harder finding is on developability. AIntibody didn't just ask whether predicted antibodies bind — it asked whether they could actually go into a patient. Solubility, stability, aggregation. Antibodies that looked right on a screen and failed in a vial. No in silico model caught those failures. What the spread across 29 teams is really showing you, the episode argues, is the field's first granular skills audit — not noise, but signal about which computational moves work and for which problems. One benchmark round, the way CASP eventually shaped structural biology over decades. Whether AIntibody achieves the same standardizing effect is genuinely unresolved. The antibody still has to stick.

Frequently asked

What did the AIntibody benchmark find about AI antibody design?

The AIntibody benchmark, published in Nature Biotechnology and organized by Specifica LLC, tested 511 computationally designed antibody sequences from 29 organizations in blinded wetlab assays. No single computational approach dominated across all tasks, and structural accuracy — even near-experimental precision on folding — did not reliably translate to binding affinity or developability.

Does AlphaFold 3 improve AI antibody design?

Inverse folding models that used AlphaFold 3-generated antibody-antigen complex structures achieved the highest average Spearman correlation and best precision@10 in the AIntibody benchmark. However, that advantage applied specifically to inverse folding as a model class — it did not mean AlphaFold 3 structures reliably predicted binding affinity or solved developability.

Why doesn't structural accuracy predict antibody binding affinity?

Binding affinity depends on molecular dynamics, solvation, and conformational entropy — factors a static structural model cannot capture. The AIntibody benchmark confirmed that an antibody can have near-perfect predicted backbone geometry and still fail to bind experimentally, because structure and affinity are genuinely distinct problems in computational drug discovery.

What is developability in antibody drug discovery, and why does it matter for AI models?

Developability refers to an antibody's solubility, stability, and low aggregation — the biophysical properties required to manufacture and administer it safely. The AIntibody benchmark evaluated developability alongside affinity and found that no computational model reliably flagged aggregation or stability failures, meaning candidates that appeared promising in silico failed in wetlab biophysical assays.

What is precision@10 in antibody benchmarking?

Precision@10 measures whether a computational model's top ten predicted binders are actually among the top experimentally confirmed binders in wetlab affinity assays. In the AIntibody benchmark, inverse folding models using AlphaFold 3 structures achieved the best precision@10 within their model class, making it one of the key metrics distinguishing computational approaches.

Grounded in 7 sources
Apodex Discovery: Reality Benchmarks and Environments ... · arxiv.org
Accelerating antibody discovery and optimization with high-throughput experimentation and machine learning · link.springer.com
nature.com · nature.com
AIntibody: an experimentally-validated in silico antibody discovery design challenge · pmc.ncbi.nlm.nih.gov
AI Antibody Design Just Got Its First Honest Report Card — And the Grade Is Complicated · clinicaltrialvanguard.com
Isomorphic Labs & AlphaFold: AI Drug Discovery in Trials | IntuitionLabs · intuitionlabs.ai
AIntibody: An experimentally-validated in silico antibody ... · ora.ox.ac.uk
Read transcript

Dr. Nathan Hayes: Maya, good to be back — I want to start with a scene, if that's okay, because I think it sets up everything.

Maya Chen: Yeah, go — I'm in.

Dr. Nathan Hayes: A computational biologist — somewhere, probably staring at two monitors — finishes running their AI model on an antibody target. They have sequences. Their model says these will bind. And now they have to do something genuinely uncomfortable: they have to submit those sequences into a blinded challenge before a single wetlab experiment is run. No calibration afterward. No taking it back.

Maya Chen: Hmm — and blinded means what, exactly? Like, blinded to what?

Dr. Nathan Hayes: Blinded to the experimental outcome. Here's the analogy that actually clicked for me: it's like being asked to write down your recipe before you taste the dish — and then someone else cooks it and grades you on how it turns out. You can't revise after seeing the result. That's what AIntibody, organized by Specifica LLC — Andrew R. M. Bradbury's group — actually enforced. Twenty-nine organizations sent in their best computational guesses, five hundred and eleven antibody sequences total.

Maya Chen: And the results came back in Nature Biotechnology.

Dr. Nathan Hayes: Right — and what came back was, to put it plainly, the field's first honest experimental report card on in silico antibody discovery. Computational methods designing sequences, wetlab measuring whether they actually bind. That gap between those two things is what the whole benchmark is about.

Maya Chen: The gap between what the model sees and what the biology does.

Dr. Nathan Hayes: Right — but that gap just got a lot more specific. Because the results didn't come back uniform. Inverse folding models — the ones that took AlphaFold 3-generated full antibody-antigen complex structures and worked backward from there — those hit the highest average Spearman correlation and the best precision@10 among their model class.

Maya Chen: So AF3 won.

Dr. Nathan Hayes: That's the headline that wants to write itself, yes. But — importantly — it won one metric class. Inverse folding, using AF3 structures. That's narrower than 'AlphaFold 3 solved antibody discovery.'

Maya Chen: And precision@10 is — just to be concrete — that's whether your top ten predicted binders are actually among the top experimental binders when you run the wetlab. Like, did your ranked list hold up against the biology.

Dr. Nathan Hayes: Exactly that. And experimental affinity is the ground truth — there's no computational substitute for that measurement. The wetlab is the arbiter.

Maya Chen: Here's the turn that I think is actually sort of — mm — destabilizing. A researcher reads this Friday afternoon. Inverse folding models, AlphaFold 3 structures, top Spearman correlation — they're feeling cautiously good. And then they keep reading and they find that structural accuracy, even near-experimental precision on how the antibody folds around the target, does not reliably translate to affinity. The antibody predicted to stick... just doesn't. In the wetlab. That's not a rounding error, that's — wait, that's the whole problem the benchmark is surfacing.

Dr. Nathan Hayes: Because structure and binding affinity are genuinely different problems. You can have a perfect structural model — sub-angstrom backbone geometry — and a completely wrong binding constant. The mechanism driving affinity involves dynamics, solvation, conformational entropy. A static structure doesn't capture those.

Maya Chen: So the field built a cultural story — structure prediction is the hard part, and AlphaFold cracked it — and the benchmark is quietly saying, no, that wasn't actually the hard part for drug discovery.

Dr. Nathan Hayes: Which is why Isomorphic Labs and Google DeepMind producing the best structural inputs still doesn't close the loop. The advance is real and measurable. The leap to 'therefore we can predict binding' — that's where the data breaks.

Maya Chen: Which actually reframes what the 29-team variation is doing. Because I think the instinct is to read that spread — no single approach dominating across 511 sequences — as the field being messy. But it's sort of, wait, actually — that's not noise. That's the first honest audit of which tools fit which problems.

Dr. Nathan Hayes: That's the read I'd push, yes. Now — there were three distinct tasks. In silico affinity maturation, starting from phase-one sequencing outputs. Affinity ranking within HCDR3 clusters — that's the most variable loop on the heavy chain, the region doing most of the binding work. And CDR design for protein targets that weren't even in the selection output. Different problems. Different winners.

Maya Chen: So less horse race, more — skills audit.

Dr. Nathan Hayes: Precisely. And that granularity is exactly what CASP produced — eventually. Critical Assessment of Structure Prediction ran for decades, multiple rounds, before it created the conditions for AlphaFold's Nobel Prize-winning advances. AIntibody is one round.

Maya Chen: Hmm — one round. So imagine a team that just nailed HCDR3 ranking but finished mid-pack on CDR design. In one round, do they even know if that's a real signal or just... the shape of this particular dataset?

Dr. Nathan Hayes: They don't. And CACHE ran into the same wall — that benchmark found no algorithm can consistently select potent drug-like binders. That's part of why AIntibody was needed. But whether AIntibody achieves the standardizing effect CASP eventually did — genuinely unresolved.

Maya Chen: Right — but the part that makes this worse, I think, comes a bit later — what the benchmark actually says about developability sitting alongside affinity, and what that means for sponsors who've already treated computation as a done deal.

Dr. Nathan Hayes: Yes — that's where the pressure lands on actual clinical timelines.

Maya Chen: Because the variation across those 29 teams isn't a problem to fix before the field matures. It's the finding. It's the field telling itself — finally, with experimental affinity as the ground truth — which computational moves actually work, and for what.

Dr. Nathan Hayes: And developability is where that gets concrete in a way that actually stops you cold. Because AIntibody didn't just test whether predicted antibodies bind. It evaluated solubility, stability, low aggregation — the full biophysical profile. An antibody that binds beautifully in silico and then clumps in a vial is not a drug. It's a failed experiment.

Maya Chen: Wait — so the benchmark was testing a harder question than I had in my head.

Dr. Nathan Hayes: Considerably harder. Can you make this thing and actually put it in a patient? Not just — does the structure look right on a screen.

Maya Chen: Which means — okay, here's the image I can't shake. Picture a research director at a biotech, late quarter, scrolling through the Nature Biotechnology paper on her tablet. She told her CMO, maybe eight months ago, that their computational platform had de-risked candidate selection. That was the claim. And now she's looking at the performance distribution across 29 organizations and she's... finding her platform in the middle. Not because affinity was wrong — maybe affinity was fine — but because developability, the aggregation numbers, the stability data, that's where the wetlab ground truth came back and said no. No computational model flagged that. None could.

Dr. Nathan Hayes: And that's not a technical conversation she's having next. That's — what did she actually validate versus what did she claim she validated.

Maya Chen: Nature Biotechnology called it a 'complicated' report card. I think 'like cold water' was the phrase. And that framing is doing real work — it's not saying the tools failed. It's saying the sponsors who treated computation as a done deal now have experimental affinity and developability data that they cannot reframe around.

Dr. Nathan Hayes: Now — no confirmed regulatory guidance has been issued. That consequence is prospective. But the benchmark makes the question newly concrete in a way it simply wasn't before this paper existed. A CMO asking 'where does our platform sit on this distribution' now has a real distribution to point at.

Maya Chen: Right — and that's sort of the thing that sticks. Wetlab validation isn't a formality you run at the end. The AIntibody design makes that structural — you cannot work backward from the result. Experimental measurement is the floor everything else is built on, and the benchmark just made that impossible to paper over.

Dr. Nathan Hayes: Inverse folding models using AlphaFold 3 structures actually moved the needle on a real metric. Highest average Spearman correlation, best precision@10 in that class. That's a starting point. Not a shortcut to the clinic. A starting point.

Maya Chen: Yeah. And AIntibody exists — I mean, the whole reason Andrew Bradbury's group built this — is that progress needs a witness. Something that can say: here is what actually happened when the antibody had to stick. Not what the model said would happen.

Dr. Nathan Hayes: The antibody still has to stick.

Maya Chen: That's — yeah. That's the whole thing, isn't it. Thanks for working through this with me.

In silico antibody design faces the reality check of binding assays · Onpode