We're publishing Equi-mRNA, a codon-level equivariant language model for mRNA, the representation underneath several of the co-scientist's design and prediction programs. The paper is peer-reviewed and published at NeurIPS 2025.
The problem with treating mRNA as text
The genetic code is degenerate: most amino acids are encoded by more than one codon. GCU, GCC, GCA and GCG all specify alanine: different sequences, identical function. A model that treats mRNA as an ordinary token sequence has to learn that equivalence from data, separately, for every motif it ever appears in. That's sample-inefficient, and it doesn't generalize: the model can be confidently right on the training distribution and wrong the moment a familiar motif shows up in an unfamiliar arrangement.
Mapping degeneracy to a symmetry group
Equi-mRNA encodes that structure directly instead of hoping the model finds it eventually. Synonymous codons are mapped to cyclic subgroups of SO(2), so codon substitutions that preserve amino-acid identity become rotations the model is equivariant to by construction. The prior is enforced two ways: an auxiliary equivariance loss during training, and symmetry-aware pooling when codon-level representations are aggregated into a sequence-level one.
The bet: geometry the architecture already knows is geometry the model doesn't have to spend parameters memorizing. What's left over is capacity for the biology we don't yet have a closed-form description of.
What the numbers show
Trained on 25M sequences drawn from 56M RefSeq entries and evaluated across six biological benchmarks, Equi-mRNA holds its own against models several times its size, and beats them:
accuracy: +~10% (avg. across 6 benchmarks)
generative fidelity: 4.3x (Fréchet BioDistance)
functional preservation: +~28%
parameters: 15M vs. 50M–82M for the compared baselines
Parameter count isn't a vanity metric here: it's a proxy for how much of a baseline's size is spent relearning symmetry the encoding gives Equi-mRNA for free.
Checking the model against real biology
The clearest evidence the model learned something real, rather than a benchmark artifact, comes from what falls out of it unsupervised. The learned codon rotations correlate with GC-content bias (r = 0.98, R² = 0.97) and with tRNA abundance patterns (ρ = −0.69). Both are independent, well-characterized biological signals the model was never directly trained to predict. When a representation's internal geometry lines up with measurements nobody asked it to reproduce, that's a stronger signal than one more point of benchmark accuracy.
Where it shows up
Equi-mRNA is the representation behind expression prediction, stability assessment, mRNA generation, and therapeutics design inside the co-scientist. The paper is on arXiv (2508.15103) and peer-reviewed at NeurIPS 2025. It's the architecture running underneath a candidate's predicted expression or stability the next time you see one inside a platform run.


