← All models
Benchmark

MIRAGE

Measuring Interpolation and Redundancy in Affinity GEneralization

A single held-out correlation cannot separate transferable binding principles from repeated exposure to protein families already dense in public structural databases. MIRAGE treats historical public family support as an explicit experimental variable, and applies the same axis to affinity and to pose.

Preprint · September 2026 · Mehdi Yazdani-Jahromi, Sanjay Padhi, Ivan Garibay · pip install mirage
Read the paper ↗GitHub ↗Hugging Face ↗
+0.45
Family gap, Nesso-1
+0.62
Family gap, Boltz-2
≈0
Family gap, every control
+0.64
Pose gap, Boltz-2 without an MSA

Affinity accuracy versus protein-family support

Matched-pK Pearson r within each support bin. Co-folders rise steeply with the number of structures their family already had; family-disjoint controls stay flat.

Read as a table ↓
0.60.40.20.0-0.112–56–2021–8081–300301+singletonProtein-family size (MMseqs2 30% cluster)Pearson rBoltz-2Nesso-1ligand-kNNRF-QSARmol. weight
Affinity co-folders Nesso-1, Boltz-2Family-disjoint controls RF-QSAR, ligand-kNNTrivial descriptor molecular weightDashed line = second method in the class

The family generalization gap

Gm is the change in matched-pK accuracy from singleton families to the best-supported ones. The two co-folders carry a large positive gap; the family-disjoint controls and the classical empirical scorer do not. ΔG is the paired difference on the co-folder's own complexes, since prediction coverage differs by method.

#Method95% CIon Gm
Boltz-2Affinity co-folder+0.62––[+0.13, +1.02]
Nesso-1Affinity co-folder+0.45––[+0.28, +0.61]
gninaDocking / deep net+0.07+0.38[+0.17, +0.59]+0.52[−0.08, +1.09][−0.14, +0.27]
RF-QSARFamily-disjoint control-0.01+0.53[+0.26, +0.75]+1.11[+0.49, +1.71][−0.17, +0.16]
ligand-kNNFamily-disjoint control-0.03+0.48[+0.25, +0.72]+0.92[+0.22, +1.57][−0.17, +0.11]
sminaDocking / deep net-0.04+0.32[+0.09, +0.59]+0.84[+0.10, +1.47][−0.18, +0.19]

Table 1. Matched-pK window, MMseqs2 30% clustering, two-level bootstrap 95% CI. Both co-folder gaps exclude zero; every family-disjoint control interval spans it. Every paired ΔG excludes zero except gnina against Boltz-2. Restricting to matched samples moves the controls' gaps more negative than their full-sample values, so an unpaired comparison understates the interaction.

The claim

Co-folding models such as Boltz-2 and Nesso-1 report protein-ligand affinity correlations near Pearson 0.6 and are proposed as fast surrogates for free-energy perturbation. That single number cannot separate a model that has learned binding physics from one that recognizes protein families it has seen many times in the PDB.

The distinction decides programs. A medicinal chemistry campaign succeeds or fails on novel targets, the ones least represented in public structural and affinity data. So MIRAGE asks the question the overall correlation hides: how does accuracy change as a target's family becomes less familiar?

The central result is an interaction, not a ranking. Frontier co-folders depend strongly on historical public family support. Family-disjoint shallow controls do not. On novel families the practical ranking inverts, and on one external low-support target neither co-folder convincingly beats molecular weight.

We call this redundancy-driven inflation rather than leakage: proprietary training sets cannot be reconstructed, so family density stays a proxy for exposure rather than evidence of it, and the design does not separate memorisation from other family-correlated difficulty.

Definitions
Family support, Sf
For a target in family f, the inclusive number of PDBbind structures in the same MMseqs2 30% sequence-identity cluster, the evaluation complex included, so a singleton family has Sf = 1. It is historical public support through 2019: every evaluation complex is a 1982–2019 PDBbind deposition. Binned into {1, 2–5, 6–20, 21–80, 81–300, 301+}.
Matched strata
Within each support bin, Pearson r is computed over a matched pK window of [4.5, 8.0], holding label spread, and therefore attainable correlation, approximately constant across bins.
Family generalization gap, Gm
Gm = r(Sf ≥ 301) − r(Sf = 1)The change in accuracy from singleton to well-supported families. A positive Gm means performance increased with family support. Estimated with a two-level bootstrap: resample families with replacement, then resample ligands within each family, so the interval reflects both family-level and ligand-level variability.
Controls
The honest counterpart to a frozen co-folder, whose test families sit inside its training window, is a model that cannot have seen the test family. A random-forest QSAR model and a ligand-kNN baseline are evaluated under 5-fold family-disjoint cross-validation (GroupKFold on family id), so every prediction is on a held-out family.
Dose response

The chart at the top of this page, as numbers. Row maxima in bold: the co-folders peak at the most redundant bin, the controls at the most novel one.

Method12–56–2021–8081–300301+
Nesso-10.1050.2250.3780.4210.4370.551
Boltz-2-0.030-0.007-0.0080.4200.3710.585
RF-QSAR0.2590.2640.2570.2540.1960.177
ligand-kNN0.2150.2230.1920.2360.1740.180
mol. weight-0.0320.009-0.1010.1070.070-0.010

Table 2. Matched-pK Pearson r by protein-family support.

The dependence survives the measured covariates. Regressing per-target absolute error on log10 Sf together with ligand similarity, protein length, publication year, within-family affinity variance, assay type, and the affinity regime, the family-support coefficient stays large and significant: Nesso-1 −0.130 (t = −5.1), Boltz-2 −0.152 (t = −3.4). Standard errors are CR1 cluster-robust, clustered on protein family (1,410 and 373 clusters), matching the dependence structure the two-level bootstrap assumes elsewhere.

Nor is the gradient exact-target repetition. Adding a repetition indicator leaves the coefficient unchanged (−0.130 → −0.135, t = −5.2) with the indicator insignificant, and restricting to targets whose sequence appears nowhere else in the corpus leaves the gap undiminished: +0.49 [+0.23, +0.74] against +0.45 unrestricted, with every control still flat. Repetition amplifies the effect rather than creating it.

Two axes

Two memorization channels are available: co-folders can exploit protein-family familiarity, and ligand models can exploit chemical similarity. Matching affinity distributions alone does not separate them, so MIRAGE stratifies by both at once.

family support \ ligand
novel ligand
mid
similar ligand
novel (1–5)
+0.37n = 170
+0.10n = 228
+0.11n = 325
mid (6–80)
+0.44n = 134
+0.48n = 253
+0.35n = 353
redundant (81+)
+0.51n = 152
+0.53n = 270
+0.48n = 262

Table 3. Nesso-1 Pearson r, matched-pK. Accuracy climbs down the family axis, from about 0.1 to about 0.5, and is flat across the ligand axis. The memorization channel for co-folders is protein-family familiarity, not ligand chemistry.

A predictor that outputs nothing but the mean pK of the target's family scores 0.564 under random splitting, and approximately zero under family-disjoint splitting by construction. That is close to the co-folders' pooled 0.588. Much of what a random split measures is recoverable protein-family identity rather than transferable modelling of protein-ligand binding. This tests one ligand statistic, nearest-neighbour ECFP4 Tanimoto: scaffold identity, ligand size and congeneric structure remain untested, so MIRAGE does not close the ligand channel.

Modality ablation

A model given only the protein, blind to the ligand and so unable in principle to know the affinity, still reaches 0.64 under random splitting. It collapses under family-disjoint splitting. Each row below is one feature set, from its random-split score to its family-disjoint score.

family-mean (identity)target identity alone
0.564≈0
–
ligand only (ECFP4)no protein
0.6760.503
−0.17
protein only (composition)blind to the ligand
0.6370.248
−0.39
protein only (ESM-2 650M)blind to the ligand
0.6530.330
−0.32
full (ligand + protein)both modalities
0.7400.456
−0.28
0.00.20.40.60.8
Random splitFamily-disjoint split

Table 4. Random forest, Pearson r. Gap column is the drop from random to family-disjoint splitting.

Novel target

The external arm asks whether the low-support result reproduces on a newly released ligand series. The target is the OpenBind EV-A71 2A protease, which has two prior public family structures, so it sits at the low-support end of the curve; what post-dates every documented cutoff is the ligand series and its affinity measurements, not the protein fold. Because no evaluated model discloses an affinity-data cutoff, this is a zero-shot temporal evaluation, not a prospective one.

Ranking 643 compounds against it, molecular weight sits second from the top. Nesso-1 exceeds it only marginally, and Boltz-2, Gnina, smina, AEV-PLIG, clogp and AQ-Affinity all fall below it.

Nesso-10.486+0.23
molecular weight0.469–
Gnina0.431-0.08
Boltz-20.395+0.11
smina0.229-0.01
AEV-PLIG0.228-0.14
clogp0.163-0.06
AQ-Affinity0.134-0.02

Table 5. Spearman ρ over the 643 compounds with complete predictions, and ρ after partialling out molecular weight (right column). Five of the seven methods are negatively informative once size is removed: their apparent ranking skill on this target was size tracking.

Two external sources agree with this reading. Nesso-1's own report states its advantage over molecular weight on this target is not statistically convincing, and the OpenBind release itself flags molecular weight as a strong baseline for the campaign. It is reported here as one external target, not population-level proof: a single congeneric series cannot carry that weight.

Docking scorers

The gradient is not a property of co-folding, or of model size. Re-scoring the crystal pose separates two regimes: a classical empirical scoring function is weak overall and shows no improvement with family support, while a CNN score trained on PDBbind is far stronger and carries the same dependence as the co-folders.

ScorerTypePooled rNovel familiesRedundant familiesSlope (t)
smina (Vina)classical empirical0.0950.1670.061+0.018 (t = +2.4)
gnina (CNNaffinity)learned0.4440.3300.423−0.060 (t = −2.9)

Table 6. Docking scoring functions on the affinity set. smina's slope is significantly positive, the sign molecular weight also shows: error rises slightly as families become better represented. We state that as measured rather than as flatness, and offer no mechanism for the sign.

smina and Vina are regression-fitted empirical functions, not free-energy calculations, which is why this contrast is not described as physics versus learned. Two scorers do not establish a law about every learned function, and the paper does not state one.

Pose prediction

The same family axis applies to geometry. Crystal ligands are redocked into 50 singleton and 50 redundant targets, scored by symmetry-corrected RMSD, with success at RMSD < 2 Å and a missing pose counted as a failure. Pose-Gm is redundant minus novel success.

EngineInputNovelRedundantPose Gm
sminaclassical empirical docking62%86%+0.23
SigmaDockneural diffusion docking52%80%+0.28
DiffDock-Lneural diffusion docking18%54%+0.36
Chai-1co-folding, no MSA28%64%+0.36
Boltz-2co-folding, no MSA2%66%+0.64
Boltz-2co-folding, +MSA46%66%+0.20

Table 7. Redocking, n = 50 per bin. The first three engines are given the crystal receptor, where smina sets a +0.23 difficulty reference for how much harder the novel targets are; the co-folders build the complex de novo from strictly less information, so their gaps are not that reference plus a memorisation excess and nothing is subtracted.

The controlled comparison for a co-folder is therefore within-model. With weights, architecture and target set fixed, a ColabFold MSA raises Boltz-2's novel-family success from 2% to 46% while redundant success stays at exactly 66%, collapsing pose-Gm from +0.64 to +0.20. The effect is one-directional: 22 of 50 novel complexes flip from failure to success and none the reverse, and both settings posed all 100 targets. Chai-1 gains on novel families too (22% to 33%) but loses on redundant ones and suffers run attrition, so it counts as partial directional replication.

Supplying evolutionary context therefore removes much of the pose deficit associated with low public family support. That deficit is localisable and at least partly a missing-input problem, not an irreducible one. Family novelty here is defined on PDBbind structural support while ColabFold retrieves from far larger sequence databases, which is precisely why the rescue is possible.

Confidence proxies

Chai-1 and ESMFold2 predict structure and an interface confidence (ipTM), not affinity. Scoring ipTM as an affinity proxy is a weaker signal than a trained affinity head, and the two proxies differ instructively.

MethodnPooled rNovel familiesRedundant familiesSlope (t)
Chai-1 (ipTM)5980.3940.1550.461−3.4
ESMFold2 (ipTM)5990.1830.1020.022−0.5

Table 8. Interface ipTM scored as an affinity proxy.

Chai-1's ipTM carries the same family-support dependence, rising from 0.155 on novel families to 0.461 on redundant ones (slope t = −3.4). Since Chai-1 has no affinity head, this is consistent with the dependence residing in the learned structural representation rather than only in a regression head fitted to affinity labels, though ipTM is a confidence score and is itself responsive to difficulty. ESMFold2's ipTM is a weak and noisy proxy with no significant dependence (t = −0.5), so the effect is not universal across confidence proxies; that arm is thin and nothing further is drawn from it.

LBA30

The thesis is not specific to MIRAGE's own dataset. ATOM3D LBA is the field's standard structure-based affinity-generalization benchmark: PDBbind-refined complexes split by 30% protein sequence identity, on which recent geometric deep nets report large gains. No published LBA30 result reports a ligand-only, molecular-weight, or QSAR control. The field compares only against other structure and sequence deep nets. Those controls are supplied below.

IPBind0.732
EHIGN0.612
GIGN0.586
3DCNN0.550
GNN0.545
DeepDTA0.472
ligand-only RF (ECFP4)0.440
molecular weight0.436
ligand-kNN0.417
ENN0.389
clogp0.202

Table 9. Pearson r on pK, full official test split (n = 490; 489 for the fingerprint-based controls). Deep-net numbers are from the literature; the three ligand-only controls in bold are ours and, to our knowledge, have not previously been reported.

Two conclusions follow, and both are stated. First, the top methods genuinely learn structural signal: IPBind exceeds the ligand-only control by roughly 0.29, which is strong evidence of predictive signal beyond a ligand descriptor. Second, and more consequential for the field, molecular weight alone scores 0.436, above ENN, near DeepDTA, and within 0.11 of 3DCNN and GNN. Small reported margins over those methods were never contextualised against a trivial descriptor, and a reader of the published leaderboard has no way to tell which entries clear it.

The same control reconciles LBA30 with MIRAGE. A 30% sequence split is a far weaker generalization test than singleton-family stratification: a ligand-only model reaches 0.440 on LBA30 but only 0.31 to 0.41 on MIRAGE novel families. So 0.73 and 0.31 do not contradict each other. They measure different stringencies, and only the second resembles the situation of a genuinely new target.

Datasets
Redundancy set
18,759

PDBbind-derived complexes with an exact Kd, Ki, or IC50 label (pK), each annotated with protein-family support Sf and ligand nearest-neighbour ECFP4 Tanimoto. A balanced 3,360-target core subset (560 per support bin) ships for quick evaluation.

Temporal set
649

Compounds with measured Kd against one low-support target: the OpenBind EV-A71 2A protease (CC0), which has two prior public family structures. The ligand series and its measurements post-date every documented cutoff; the protein fold does not. 649 compounds ship in the set; 643 carry predictions from every method and are the ones scored in the ranking tables. Molecular weight, clogp, Gnina, smina, AEV-PLIG, and AQ-Affinity baselines ship with it.

Reporting protocol

What the paper recommends anyone reporting affinity accuracy should do:

  • Weight families equally, not complexes
  • Report mean within-stratum correlation, not pooled-across-target correlation
  • Report Pearson, Spearman, centred RMSE, and pairwise ranking accuracy
  • Use a two-level bootstrap: families first, then ligands within families
  • Show sensitivity to the clustering choice
  • Ship a baseline suite: family-mean (target identity), molecular weight, clogp, ligand-kNN, and RF-QSAR
  • Report the family gap and the novel-family number, not a single overall correlation
Limitations
  • The redundancy set is pre-cutoff PDBbind, and Sf is a within-corpus support proxy, not the models' exact training counts
  • The temporal set is a single target with a congeneric series, and is zero-shot rather than prospective: no evaluated model discloses an affinity-data cutoff
  • Sf is historical public family density, a proxy for exposure rather than evidence of training exposure, and the design does not separate memorisation from other family-correlated difficulty
  • Within-exact-target ranking is not estimable on this corpus: requiring five ligands per identical sequence leaves six targets for Nesso-1 in the redundant stratum and none in the novel one
  • No arm varies training-set redundancy, so the MSA result localises the pose deficit without establishing how family redundancy acts during training
  • Frozen public weights are evaluated; no model is retrained on de-leaked data
  • Family support at 30% identity is one axis. Pfam-based and pocket-based clustering are natural sensitivity analyses
  • The confidence-proxy results use ipTM, not a trained affinity head
  • AlphaFold3 has open code but request-only weights and no affinity head. It ships as a pluggable stub and is not evaluated
Reproduce it

MIRAGE is an installable package with a Hugging Face dataset. Scoring a model means wrapping it in one function.

from mirage import run_redundancy, run_temporal

def predict(sequence: str, smiles: str) -> float:
    ...                      # your model, higher is stronger

run_redundancy(predict)      # family-support curve + G_m
run_temporal(predict)        # external low-support target

Function names follow the paper. The repository documents the current CLI flags and dataset configuration names.

A CLI scores a predictions CSV instead. The public release includes all metrics (Gm, matched-pK, the multivariate regression, and the two-level bootstrap), every baseline, the dataset pipeline, and the released predictions used to reproduce each number on this page.

FAQ

How protein-ligand affinity and pose accuracy depend on historical public protein-family support: the number of PDBbind structures through 2019 in the same MMseqs2 30% sequence-identity cluster. Standard benchmarks report one overall correlation, which averages over that variable. MIRAGE treats it as the independent variable and reports accuracy across the whole curve, for affinity and for pose.

Score your model against the whole curve.

One overall correlation hides the only regime that matters for a new target. MIRAGE reports both.

View on GitHub ↗