Ben Okonkwo: Marcus, hey — you look like you've been stewing on something.
Marcus Vale: I have. I keep thinking about a scene — there's a structural biologist, she's at her workstation at 2am, it's late 2020, CASP14 results are coming in. And she's watching AlphaFold 2's predictions render next to the crystal structures she probably spent years collecting. And they match.
Ben Okonkwo: Hm — near-experimental accuracy. From sequence alone.
Marcus Vale: Demis Hassabis, John Jumper, the DeepMind team — they basically told fifty years of structural biology: we can do in seconds what you spent careers on. And CASP14 is where that became undeniable.
Ben Okonkwo: Now — the reason it could do that is the Protein Data Bank. The PDB. Decades of protein structures deposited by experimentalists, and AlphaFold 2 trained on all of it. But here's what I think people missed in that 2am moment of wonder — what exactly is in the PDB?
Marcus Vale: Static structures. Proteins sitting still.
Ben Okonkwo: Mostly ligand-free, mostly crystallized — frozen in one position. So AlphaFold 2 learned an extraordinarily precise answer to one question: given this sequence, what's the stable resting shape? One structure. One state.
Marcus Vale: The sculptor analogy. The model is like a sculptor who has studied thousands of photographs of a person's face — and can now carve a perfect likeness from memory. But every photograph was taken while that person was completely still.
Ben Okonkwo: And you ask the sculptor: what does her face look like when she laughs? Or when she's startled? The photographs never showed that.
Marcus Vale: Which maps directly onto — what does this protein look like when a drug binds to it and forces it into a different shape?
Ben Okonkwo: Exactly the problem. And what we're really trying to work out today is: is that a gap you can engineer around, or is it a hard limit baked into what AlphaFold 2 was ever designed to do?
Marcus Vale: Because the answer changes everything for drug discovery.
Ben Okonkwo: And drug discovery is exactly where that limit bites hardest. So picture the same biologist, six months later — she's handing her AlphaFold 2 prediction to a drug discovery team. They're targeting Abelson kinase. And the first thing they ask is: okay, great, we have the resting shape. What does this protein look like when our compound actually binds?
Marcus Vale: Wait — so the version of the protein the drug actually encounters is a completely different shape?
Ben Okonkwo: Different enough to matter, yes. That's induced fit — the protein physically reshapes its three-dimensional structure when the ligand arrives. Not a small tweak. This is core to how enzymes catalyze reactions, how receptors signal. It's not a rare edge case, it's the mechanism.
Marcus Vale: And AlphaFold 2 never saw a ligand. During training. At all.
Ben Okonkwo: Right — no ligands, no ions, no solvent, no post-translational modifications during prediction. The model has no physical inputs that would let it anticipate what happens when a small molecule lands. So it gives you the apo state — the ligand-free form — and the holo state, the bound form, can be substantially different. With Abelson kinase, that gap is measurable and it's consequential for whether you're designing a drug against the right target shape.
Marcus Vale: So the drug discovery team is — basically optimizing against the wrong conformation.
Ben Okonkwo: Potentially, yeah. And then there's a second layer on top of that — allostery. Binding somewhere else on the protein, at a completely separate site, changes the shape over there. Activity at a distant location is altered. AlphaFold cannot see that either, because it doesn't know a ligand is present anywhere.
Marcus Vale: Hold on. So you have induced fit, which is local reshaping at the binding site. And then allostery, which is — the signal travels across the protein and reshapes something far away. And AlphaFold 2 is blind to both.
Ben Okonkwo: Blind to both. Now — here's the part I find genuinely strange when you sit with it. The reason AlphaFold works at all is protein folding cooperativity. Most proteins snap sharply between two well-defined states, folded and unfolded, without populating many intermediates. Two-state kinetics. That's what makes the resting structure learnable — there's a single tractable answer.
Marcus Vale: The same property that makes static prediction tractable is what makes the transition invisible.
Ben Okonkwo: Exactly — the cooperativity that gives you a clean ground truth for the apo state also means the protein resists intermediate conformations. So the transition state, the shape the protein moves through going from apo to holo — it's thermodynamically unstable, it doesn't sit still long enough to appear in the Protein Data Bank, and AlphaFold never learned it.
Marcus Vale: Okay, so this isn't a training data volume problem. It's architectural. You can't learn what was never stable enough to be deposited.
Ben Okonkwo: That's the distinction that matters. And it's why NeuralPLexer — published in Nature Methods, April 2024 — is interesting. It was built to predict both the apo and holo forms, with the ligand encoded as an actual input. It outperformed AlphaFold 2 on binding-induced conformational change. But the question I can't answer yet is whether that generalizes — allosteric proteins, different protein families — or whether it works on a narrow slice.
Marcus Vale: And the market is not waiting for that answer. Drug teams are already using AlphaFold 2 output as if the apo state is close enough. That's the bet being made right now, whether anyone admits it or not.
Ben Okonkwo: And that bet has a direct cost that isn't abstract. If you're using the apo conformation to predict binding affinity — how tightly your drug candidate grips the target — you're feeding the wrong shape into your calculation. The number you get back is built on a structure the protein doesn't actually adopt when the drug is present.
Marcus Vale: So the affinity number is just... wrong.
Ben Okonkwo: Potentially wrong in a way that looks precise. You run the calculation, you get a binding energy, it has decimal places. The pipeline moves forward. But if the holo conformation differs significantly from apo — and on proteins like Abelson kinase it does — then the pocket geometry you were docking into doesn't exist when the drug actually arrives.
Marcus Vale: Confident decimal places on the wrong shape. That's — yeah, that's the part that should scare people.
Ben Okonkwo: Now picture a medicinal chemist — she's optimizing a lead compound, iterating structures week by week, and every affinity prediction she's trusting has AlphaFold 2's apo output underneath it. She doesn't know the binding site reshapes on contact. She's chasing a conformation the protein sheds the moment her molecule walks in.
Marcus Vale: Which is exactly what NeuralPLexer was built to fix, right? April 12, 2024, Nature Methods — it encodes the ligand as an actual input and predicts both the apo and holo forms. That's the architectural move AlphaFold 2 never made.
Ben Okonkwo: That's the claim. And on proteins with large binding-induced conformational changes, NeuralPLexer outperforms AlphaFold 2 — the benchmark supports that. But I want to be careful about the step from 'outperforms on a benchmark' to 'solves ligand binding affinity prediction for drug discovery.' Those are not the same sentence.
Marcus Vale: What's the gap?
Ben Okonkwo: Generalization. Does NeuralPLexer work on allosteric proteins — where the drug binds remotely and the relevant conformational change is at a distant site? Or does it excel specifically on tight-binding enzymes with obvious binding pockets? That distinction determines whether this is a tool for a slice of the problem or the whole problem.
Marcus Vale: Okay, so what about AlphaFold 3? DeepMind ships the differentiable simulation framing — multi-scale transformers, biologically informed cross-attention, geometry-aware optimization. That sounds like it's moving toward dynamic modeling.
Ben Okonkwo: Moving toward is doing a lot of work in that sentence. AlphaFold 3 still does not fully model ligands, solvent, ions, or post-translational modifications during inference. Those are precisely the inputs you need to predict context-dependent conformational outcomes. The architecture is more sophisticated, but the physical inputs are still missing.
Marcus Vale: So 'differentiable simulation paradigm' is — what, a better description of the aspiration?
Ben Okonkwo: It's a real architectural direction, I don't want to dismiss it. But show me the benchmark — specifically on proteins where binding causes more than five Angstroms of RMSD shift. Where does AlphaFold 3 land versus NeuralPLexer on that slice? I haven't seen that number.
Marcus Vale: Basically, 'moving toward' isn't shipped. DeepMind built the foundation. They haven't built what drug discovery actually needs.
Ben Okonkwo: And there's a layer underneath that that makes this harder than it looks — we'll get to it, but the real problem isn't just predicting one holo structure. It's that the ground truth is actually a whole distribution of conformations varying with pH, crowding, modifications. A single output model may be wrong in principle, not just in practice.
Marcus Vale: So are we watching a solution converge, or just getting better at describing where the wall is?
Ben Okonkwo: But 'describing where the wall is' undersells what the wall actually is. Because the wall isn't 'we need more data on holo states.' The wall is — the correct answer is not a structure. It's a distribution. Think about what that means architecturally.
Marcus Vale: A distribution of what, exactly?
Ben Okonkwo: Of conformations. The conformational ensemble — that's the actual ground truth. A protein at physiological pH, with specific ions present, with post-translational modifications on certain residues, crowded by other molecules in the cell — it doesn't sit in one shape. It samples a whole landscape of shapes. That landscape shifts when you change the pH. It shifts when a binding partner arrives. The 'true structure' is actually a conditional distribution, and it's conditional on everything.
Marcus Vale: So a model that outputs one structure — any model, not just AlphaFold 2 — is answering the wrong question by design.
Ben Okonkwo: That's the general principle, yeah. And it's not a protein-specific problem. Imagine a weather forecaster trained only on photographs of clear-sky days — not asked to name tomorrow's cloud pattern, but to give you the full probability distribution of possible weather states. The output format is just... wrong for the task. A point estimate cannot approximate a distribution. It answers a different question.
Marcus Vale: Huh. So AlphaFold's limitation is actually an instance of that — a universal ML failure mode when the target is a distribution and you've shipped a point estimator.
Ben Okonkwo: Exactly — and you can't fix it by adding parameters to the same architecture. Now, the field is actually moving toward something called conformational ensemble prediction, where conformation is a query. You condition the model on ligand, on pH, on modification state. But that's — I mean, that requires rethinking what a model should output at a foundational level. That's not AlphaFold 3 with a bigger transformer.
Marcus Vale: Okay, so who owns that? Because if the conditional ensemble model is the actual moat — not the base structure predictor — then DeepMind built the thing that got commoditized first.
Ben Okonkwo: That's — right, but the specific fact that stops me here is cooperativity. The same two-state kinetics that made AlphaFold trainable means the protein actively resists the intermediate states you'd need to sample to build that ensemble. The distribution you're trying to capture is thermodynamically thin in the places that matter most for drug binding.
Marcus Vale: So the model that owns drug discovery for the next decade isn't a structure predictor at all. It's something that treats structure as an output of context — ligand present, pH this, modification there — and returns a range.
Ben Okonkwo: And nobody has built that yet. NeuralPLexer gets you one step closer — encoding the ligand as an actual input is the right move — but it still outputs specific conformations, not an ensemble distribution. The gap between those two things is not incremental.
Marcus Vale: So is this an engineering problem, or a fundamental mismatch between output format and target that no one can resolve without a different conception of what the model is even for?
Ben Okonkwo: Both, probably. I mean — the output format mismatch is fundamental. But the deeper thing is what you'd actually have to ask. The question that matters for drug discovery isn't 'what does this protein look like?' It's... what does it look like given this ligand, this pH, this crowding, these post-translational modifications, in this cell type, under this stress condition? That's not a harder version of the same question. That's a different question entirely.
Marcus Vale: Conformational state as a query.
Ben Okonkwo: Right — and nobody's built that. AlphaFold 2 can't answer it. AlphaFold 3 isn't there yet. NeuralPLexer gets you one input closer. But the model that actually answers that question — it outputs a distribution, conditioned on everything. That's not a bigger transformer. That's a different conception of what the model is for.
Marcus Vale: I keep thinking about her. That biologist from 2am, CASP14 — watching AlphaFold 2 nail the resting shape. And now, same workstation, same researcher, but the question on her screen isn't 'what shape is this protein.' It's 'what shape will this protein be when my compound lands on it, at pH 6.8, in a hypoxic tumor cell.' AlphaFold cannot answer that. Maybe nothing can yet.