Home/Research

Introducing Equi-mRNA: the first codon-level equivariant mRNA language model

Most sequence models treat the genetic code as text and relearn its geometry from scratch. Equi-mRNA builds that geometry in: a 15M-parameter model that outperforms baselines 3-5x its size, peer-reviewed at NeurIPS 2025.

RResearch teamDeepBio ScientificAug 4, 2026 · 8 min read

We're publishing Equi-mRNA, a codon-level equivariant language model for mRNA, the representation underneath several of the co-scientist's design and prediction programs. The paper is peer-reviewed and published at NeurIPS 2025.

The problem with treating mRNA as text

The genetic code is degenerate: most amino acids are encoded by more than one codon. GCU, GCC, GCA and GCG all specify alanine: different sequences, identical function. A model that treats mRNA as an ordinary token sequence has to learn that equivalence from data, separately, for every motif it ever appears in. That's sample-inefficient, and it doesn't generalize: the model can be confidently right on the training distribution and wrong the moment a familiar motif shows up in an unfamiliar arrangement.

Mapping degeneracy to a symmetry group

Equi-mRNA encodes that structure directly instead of hoping the model finds it eventually. Synonymous codons are mapped to cyclic subgroups of SO(2), so codon substitutions that preserve amino-acid identity become rotations the model is equivariant to by construction. The prior is enforced two ways: an auxiliary equivariance loss during training, and symmetry-aware pooling when codon-level representations are aggregated into a sequence-level one.

The bet: geometry the architecture already knows is geometry the model doesn't have to spend parameters memorizing. What's left over is capacity for the biology we don't yet have a closed-form description of.

What the numbers show

Trained on 25M sequences drawn from 56M RefSeq entries and evaluated across six biological benchmarks, Equi-mRNA holds its own against models several times its size, and beats them:

accuracy:                +~10%  (avg. across 6 benchmarks)
generative fidelity:      4.3x  (Fréchet BioDistance)
functional preservation: +~28%
parameters:                15M  vs. 50M–82M for the compared baselines

Parameter count isn't a vanity metric here: it's a proxy for how much of a baseline's size is spent relearning symmetry the encoding gives Equi-mRNA for free.

Checking the model against real biology

The clearest evidence the model learned something real, rather than a benchmark artifact, comes from what falls out of it unsupervised. The learned codon rotations correlate with GC-content bias (r = 0.98, R² = 0.97) and with tRNA abundance patterns (ρ = −0.69). Both are independent, well-characterized biological signals the model was never directly trained to predict. When a representation's internal geometry lines up with measurements nobody asked it to reproduce, that's a stronger signal than one more point of benchmark accuracy.

Where it shows up

Equi-mRNA is the representation behind expression prediction, stability assessment, mRNA generation, and therapeutics design inside the co-scientist. The paper is on arXiv (2508.15103) and peer-reviewed at NeurIPS 2025. It's the architecture running underneath a candidate's predicted expression or stability the next time you see one inside a platform run.

Keep reading

All posts →
Product

Early access to Helixir opens to research teams

Sep 11, 2026 / 1 min read
Research

MIRAGE: a benchmark for protein-family generalization in affinity prediction

Sep 6, 2026 / 1 min read
Research

A random forest beat two frontier co-folders. That is the benchmark talking.

Sep 6, 2026 / 5 min read