Dr. Nathan Hayes: Maya, good to be back — I want to start with a scene, if that's okay, because I think it sets up everything.
Maya Chen: Yeah, go — I'm in.
Dr. Nathan Hayes: A computational biologist — somewhere, probably staring at two monitors — finishes running their AI model on an antibody target. They have sequences. Their model says these will bind. And now they have to do something genuinely uncomfortable: they have to submit those sequences into a blinded challenge before a single wetlab experiment is run. No calibration afterward. No taking it back.
Maya Chen: Hmm — and blinded means what, exactly? Like, blinded to what?
Dr. Nathan Hayes: Blinded to the experimental outcome. Here's the analogy that actually clicked for me: it's like being asked to write down your recipe before you taste the dish — and then someone else cooks it and grades you on how it turns out. You can't revise after seeing the result. That's what AIntibody, organized by Specifica LLC — Andrew R. M. Bradbury's group — actually enforced. Twenty-nine organizations sent in their best computational guesses, five hundred and eleven antibody sequences total.
Maya Chen: And the results came back in Nature Biotechnology.
Dr. Nathan Hayes: Right — and what came back was, to put it plainly, the field's first honest experimental report card on in silico antibody discovery. Computational methods designing sequences, wetlab measuring whether they actually bind. That gap between those two things is what the whole benchmark is about.
Maya Chen: The gap between what the model sees and what the biology does.
Dr. Nathan Hayes: Right — but that gap just got a lot more specific. Because the results didn't come back uniform. Inverse folding models — the ones that took AlphaFold 3-generated full antibody-antigen complex structures and worked backward from there — those hit the highest average Spearman correlation and the best precision@10 among their model class.
Dr. Nathan Hayes: That's the headline that wants to write itself, yes. But — importantly — it won one metric class. Inverse folding, using AF3 structures. That's narrower than 'AlphaFold 3 solved antibody discovery.'
Maya Chen: And precision@10 is — just to be concrete — that's whether your top ten predicted binders are actually among the top experimental binders when you run the wetlab. Like, did your ranked list hold up against the biology.
Dr. Nathan Hayes: Exactly that. And experimental affinity is the ground truth — there's no computational substitute for that measurement. The wetlab is the arbiter.
Maya Chen: Here's the turn that I think is actually sort of — mm — destabilizing. A researcher reads this Friday afternoon. Inverse folding models, AlphaFold 3 structures, top Spearman correlation — they're feeling cautiously good. And then they keep reading and they find that structural accuracy, even near-experimental precision on how the antibody folds around the target, does not reliably translate to affinity. The antibody predicted to stick... just doesn't. In the wetlab. That's not a rounding error, that's — wait, that's the whole problem the benchmark is surfacing.
Dr. Nathan Hayes: Because structure and binding affinity are genuinely different problems. You can have a perfect structural model — sub-angstrom backbone geometry — and a completely wrong binding constant. The mechanism driving affinity involves dynamics, solvation, conformational entropy. A static structure doesn't capture those.
Maya Chen: So the field built a cultural story — structure prediction is the hard part, and AlphaFold cracked it — and the benchmark is quietly saying, no, that wasn't actually the hard part for drug discovery.
Dr. Nathan Hayes: Which is why Isomorphic Labs and Google DeepMind producing the best structural inputs still doesn't close the loop. The advance is real and measurable. The leap to 'therefore we can predict binding' — that's where the data breaks.
Maya Chen: Which actually reframes what the 29-team variation is doing. Because I think the instinct is to read that spread — no single approach dominating across 511 sequences — as the field being messy. But it's sort of, wait, actually — that's not noise. That's the first honest audit of which tools fit which problems.
Dr. Nathan Hayes: That's the read I'd push, yes. Now — there were three distinct tasks. In silico affinity maturation, starting from phase-one sequencing outputs. Affinity ranking within HCDR3 clusters — that's the most variable loop on the heavy chain, the region doing most of the binding work. And CDR design for protein targets that weren't even in the selection output. Different problems. Different winners.
Maya Chen: So less horse race, more — skills audit.
Dr. Nathan Hayes: Precisely. And that granularity is exactly what CASP produced — eventually. Critical Assessment of Structure Prediction ran for decades, multiple rounds, before it created the conditions for AlphaFold's Nobel Prize-winning advances. AIntibody is one round.
Maya Chen: Hmm — one round. So imagine a team that just nailed HCDR3 ranking but finished mid-pack on CDR design. In one round, do they even know if that's a real signal or just... the shape of this particular dataset?
Dr. Nathan Hayes: They don't. And CACHE ran into the same wall — that benchmark found no algorithm can consistently select potent drug-like binders. That's part of why AIntibody was needed. But whether AIntibody achieves the standardizing effect CASP eventually did — genuinely unresolved.
Maya Chen: Right — but the part that makes this worse, I think, comes a bit later — what the benchmark actually says about developability sitting alongside affinity, and what that means for sponsors who've already treated computation as a done deal.
Dr. Nathan Hayes: Yes — that's where the pressure lands on actual clinical timelines.
Maya Chen: Because the variation across those 29 teams isn't a problem to fix before the field matures. It's the finding. It's the field telling itself — finally, with experimental affinity as the ground truth — which computational moves actually work, and for what.
Dr. Nathan Hayes: And developability is where that gets concrete in a way that actually stops you cold. Because AIntibody didn't just test whether predicted antibodies bind. It evaluated solubility, stability, low aggregation — the full biophysical profile. An antibody that binds beautifully in silico and then clumps in a vial is not a drug. It's a failed experiment.
Maya Chen: Wait — so the benchmark was testing a harder question than I had in my head.
Dr. Nathan Hayes: Considerably harder. Can you make this thing and actually put it in a patient? Not just — does the structure look right on a screen.
Maya Chen: Which means — okay, here's the image I can't shake. Picture a research director at a biotech, late quarter, scrolling through the Nature Biotechnology paper on her tablet. She told her CMO, maybe eight months ago, that their computational platform had de-risked candidate selection. That was the claim. And now she's looking at the performance distribution across 29 organizations and she's... finding her platform in the middle. Not because affinity was wrong — maybe affinity was fine — but because developability, the aggregation numbers, the stability data, that's where the wetlab ground truth came back and said no. No computational model flagged that. None could.
Dr. Nathan Hayes: And that's not a technical conversation she's having next. That's — what did she actually validate versus what did she claim she validated.
Maya Chen: Nature Biotechnology called it a 'complicated' report card. I think 'like cold water' was the phrase. And that framing is doing real work — it's not saying the tools failed. It's saying the sponsors who treated computation as a done deal now have experimental affinity and developability data that they cannot reframe around.
Dr. Nathan Hayes: Now — no confirmed regulatory guidance has been issued. That consequence is prospective. But the benchmark makes the question newly concrete in a way it simply wasn't before this paper existed. A CMO asking 'where does our platform sit on this distribution' now has a real distribution to point at.
Maya Chen: Right — and that's sort of the thing that sticks. Wetlab validation isn't a formality you run at the end. The AIntibody design makes that structural — you cannot work backward from the result. Experimental measurement is the floor everything else is built on, and the benchmark just made that impossible to paper over.
Dr. Nathan Hayes: Inverse folding models using AlphaFold 3 structures actually moved the needle on a real metric. Highest average Spearman correlation, best precision@10 in that class. That's a starting point. Not a shortcut to the clinic. A starting point.
Maya Chen: Yeah. And AIntibody exists — I mean, the whole reason Andrew Bradbury's group built this — is that progress needs a witness. Something that can say: here is what actually happened when the antibody had to stick. Not what the model said would happen.
Dr. Nathan Hayes: The antibody still has to stick.
Maya Chen: That's — yeah. That's the whole thing, isn't it. Thanks for working through this with me.