Home/Research

A random forest beat two frontier co-folders. That is the benchmark talking.

Affinity models report one held-out correlation near 0.6. Split that number by how many times a protein family already appears in the PDB and it comes apart. MIRAGE measures the split, and finds part of the deficit is a missing input rather than a hard limit.

MMehdi Yazdani-JahromiDeepBio ScientificSep 6, 2026 · 5 min read

Here is a result we did not enjoy producing. On protein families with five or fewer prior structures, a random forest that never sees a structure, trained on fingerprints and amino-acid composition, and explicitly held out from the family it is tested on, scores a Pearson r of 0.411. Nesso-1, a frontier co-folding model, scores 0.324 on the same complexes. The paired difference is +0.087, and the interval excludes zero.

That is not a claim that shallow models are better than co-folders. It is a claim about what the headline number on an affinity benchmark is actually measuring.

The number that hides the question

Modern co-folding models report one held-out correlation, somewhere near 0.6, and that number is doing two jobs at once. It covers targets whose protein family appears in the PDB hundreds of times, and it covers targets whose family appears once. A drug programme only ever lives in the second case. Averaging the two together is how a benchmark can look solved while the regime you care about goes unmeasured.

So we stopped averaging. MIRAGE treats historical public family support as the independent variable: for each target, count the PDBbind structures through 2019 that sit in the same MMseqs2 30% sequence-identity cluster, bin that count, and report accuracy across the whole curve instead of collapsing it to a point.

The shape is the finding

Within a matched pK window, so that wider label ranges in crowded families cannot do the work, the curve separates cleanly:

family support Sf:        1     2-5    6-20   21-80  81-300  301+
Nesso-1                 0.105  0.225  0.378  0.421  0.437   0.551
Boltz-2                -0.030 -0.007 -0.008  0.420  0.371   0.585
RF-QSAR (family-disjoint)  0.259  0.264  0.257  0.254  0.196   0.177
ligand-kNN (family-disjoint) 0.215 0.223 0.192  0.236  0.174   0.180

The co-folders climb across nearly three orders of magnitude of family support. The family-disjoint controls, which by construction cannot recognise the test family, are flat or slightly decreasing. We call that difference the family generalization gap, Gm = r(Sf ≥ 301) - r(Sf = 1). It is +0.45 for Nesso-1 and +0.62 for Boltz-2, and approximately zero for every control.

The controls are not weaker co-folders. They respond differently to familiarity. A model with nothing to recognise should be flat, and they are.

Boltz-2 is worth reading slowly: it is essentially uncorrelated with truth for every family holding fewer than 21 structures, and only becomes a useful predictor once the family is well represented.

The obvious objections, and what happened to them

One curve proves nothing on its own, so the claim rests on the conjunction of several controls rather than any single one.

  • Is it a confound? In a per-target regression of absolute error on log10 Sf, controlling for ligand similarity, protein length, publication year, within-family affinity variance, assay type and affinity regime, the family-support coefficient stays large and significant: -0.130 (t = -5.1) for Nesso-1, -0.152 (t = -3.4) for Boltz-2, with errors clustered on protein family.
  • Is it just the same target appearing twice? No. Restricting to targets whose sequence appears nowhere else in the corpus leaves the gap undiminished, +0.49 against +0.45. Repetition amplifies the effect rather than creating it.
  • Is it ligand chemistry? Stratifying by family support and ligand similarity at once, the gradient runs down the family axis and is flat across the ligand axis. The memorisation channel here is the protein.
  • Is it a handful of crowded families? Equalising the seven families in the top bin moves Nesso-1's endpoint from 0.551 to 0.541. Down-weighting heavily sampled families strengthens the effect rather than dissolving it.

And a diagnostic that should worry anyone choosing a checkpoint on a random split: a random forest given only the protein, blind to the ligand and therefore unable in principle to know which of several ligands binds more tightly, reaches r = 0.637 under random splitting. Under family-disjoint splitting it collapses to 0.248. Much of what a random split rewards is recoverable family identity.

One external target, and a trivial descriptor

The redundancy arm is retrospective, so we added an external one: a newly released ligand series against a low-support target, the OpenBind EV-A71 2A protease. Across 643 compounds, molecular weight alone reaches a Spearman ρ of 0.469. Nesso-1 reaches 0.486, a margin its own report describes as not statistically convincing. Boltz-2, Gnina, smina, AEV-PLIG, clogp and AQ-Affinity all fall below molecular weight, and five of the seven are negatively informative once size is partialled out.

That is one target with one congeneric series. It is corroborative, not population-level evidence, and we say so on the page and in the paper. The same missing control shows up on the field's standard generalization benchmark: no published ATOM3D LBA30 result reports a ligand-only or molecular-weight baseline, and when we supply one, molecular weight scores 0.436, above ENN and near DeepDTA.

The constructive half

If the story ended there it would be a complaint. It does not. The pose arm gives the cleanest controlled intervention in the study: with weights, architecture and target set held fixed, supplying Boltz-2 with a ColabFold MSA raises novel-family pose success from 2% to 46%, while redundant-family success stays at exactly 66%. The pose gap collapses from +0.64 to +0.20. The effect is one-directional: 22 of 50 novel complexes flip from failure to success, and none go the other way.

So a large part of the novel-family deficit is not an irreducible limit. It is a missing input. MIRAGE identifies not only that single-sequence co-folders lean on family familiarity, but what they are missing when they do not have it.

What we do not claim

We avoid the word leakage. Leakage means related families demonstrably crossing a claimed train and test boundary, and proprietary training sets cannot be reconstructed to show that. What we establish is narrower and still decisive: accuracy scales with public family support after the measured confounds are controlled, and matched family-disjoint models do not share that dependence. Family density stays a proxy for exposure, not evidence of it, and the design does not separate memorisation from other family-correlated difficulty.

Co-folders remain the most accurate methods available on well-supported families. Nothing here contradicts that.

What to do with it on Monday

Before trusting a learned scorer on a target, compute its family support. If the family is sparse, the number that applies is the low-support value, roughly 0.1 to 0.3, not the pooled correlation. Run a family-disjoint QSAR baseline alongside it: it costs seconds, and in that regime it wins. Validate on family-disjoint splits during development rather than at the end, because architectures and checkpoints chosen on a random split are partly chosen for recognition.

And when reporting a learned method, report its excess over a reference that does not improve with family support. That costs nothing and says more than a pooled correlation. A model that beats the reference on familiar families and merely matches it on unfamiliar ones has not earned a prospective role. One that beats it on singletons has.

The benchmark, the dataset and the released predictions are packaged as an installable harness. The full write-up, every table and the limitations are on the MIRAGE model page.

Keep reading

All posts →
Product

Early access to Helixir opens to research teams

Sep 11, 2026 / 1 min read
Research

MIRAGE: a benchmark for protein-family generalization in affinity prediction

Sep 6, 2026 / 1 min read
Research

OIL: an open codon-resolved nucleotide framework, presented at MoML 2026

Sep 1, 2026 / 1 min read