← Platform
Now live in the platform

Every recommendation now stands on one graph.

We've fused ninety-plus sources into a single evidence-typed knowledge graph: the central dogma materialised end to end, gene-level coverage across 20k+ organisms, and layers of binding, regulation, clinical, and target-discovery evidence the co-scientist reasons over. It's what turns a recommendation into something you can trace, not just trust. Counts below are from the current graph snapshot (September 2026); an entity is one canonicalised node and a relation is one directed, evidence-typed edge.

Talk to us →See it in the platform
193M+
Entities · 26 types
484M+
Relations · 80 predicates
20k+
Organisms · gene-level
123M+
Genes · across 20k+ organisms
90+
Integrated sources
Built in

We show you what doesn't work, too.

Most graphs only remember hits. This one keeps curated negatives right alongside the positives, so a "no" carries as much weight as a "yes." None of it is sampled or inferred.

30M+ positive
94M+ negative

Negatives are curated, experimentally-confirmed non-interactions and inactive screening outcomes. An inactive result records inactivity under the reported assay conditions; it does not establish non-binding in general. Positive evidence spans clinical genetics, drug–target pharmacology, precision-medicine associations, and disease ontology.

Coverage

Twenty-six entity types, one graph.

Genes dominate: 123M+ across vertebrates, plants, fungi, invertebrates, protists, and ~20k bacterial genomes. But that's just the start. Transcripts, exons, chemicals, diseases, pathways, and 20 more entity types all live in the same graph, ready to traverse together. Bars are log-scaled to keep every type visible.

gene123M+
transcript47M+
exon10M+
chemical5.6M+
organism1.2M+
variant940k+
splice_event710k+
clinical_trial590k+
protein570k+
reaction420k+
utr340k+
disease160k+
pathway130k+
protein_domain41k+
phenotype30k+
biological_process28k+
sequence_cluster27k+
gene_family21k+
anatomy14k+
molecular_function11k+
cellular_component4.1k+
protein_complex2.6k+
modification2.3k+
cell_line1.2k+
keyword1.1k+
cell_type332
Reach

From 7 species to 20k+ organisms.

The central dogma is now materialized end to end (gene → transcript → exon → coding → protein) for 20k+ organisms, with genomic-neighbor edges giving every gene its operon/synteny context. Function annotation alone now spans 13k+ species.

139
Vertebrate reference genomes
Multi-kingdom
Plants · fungi · invertebrates · protists
~20k
Bacterial genomes
121M+
Bacterial genes

The central dogma, end to endgene → transcript → exon → protein

123M+
gene
has_transcript
has_transcript →
47M+
transcript
has_exon · biotype
has_exon →
10M+
exon
eukaryotes
→ coding →
570k+
protein
reference proteome · 14k+ sp.

Genome assemblies across the tree of life were parsed into the graph, with genes created wherever a clade isn't already covered. Function annotations carry evidence-tagged relationships (enables, involved_in, located_in) across 13k+ species.

Discovery

Find everything like it, through one shared node.

Every protein connects to proxy nodes for what it shares: its family, the ligands it binds, its modifications, its functional keywords, its domains. Two proteins in the same family, or binding the same cofactor, meet at that shared node: two edges, one traversal, no custom query required.

Six shared-attribute meshes on 570k+ proteinsreference proteome

has_domain (domain architecture)2.6M+
has_keyword (function / location category)3.3M+
1.1k+ keyword nodes (e.g. 99.6k+ proteins share "Nucleotide-binding," 83.4k+ "ATP-binding," 82.1k+ "Transmembrane.")
belongs_to_family (super / sub)510k+
9.7k+ family proxy nodes: 4.4k+ proteins connect through the "protein kinase superfamily" node alone; proteins link at every level (superfamily · family · subfamily).
binds_ligand (cofactor)300k+
1k+ ligand nodes: every ATP-, metal-, heme-, NAD-binding protein connects to the shared ligand.
has_ptm (modification type)250k+
354 PTM nodes (phosphoserine, N-glycosylation…), alongside the function edges that already connect proteins by shared activity.
Depth

From binding to bedside, modeled.

The molecular, clinical, and regulatory core below sits alongside the tree-of-life, shared-attribute, and RNA layers above; 80 predicates in all, each one a relationship the co-scientist can reason across. Bars are scaled within each layer.

Drug–target binding106M+

binds · positive12M+
binds · negative94M+
mechanism_of_action12k+

ADME / PK & CYP64k+ edges + 92k+ ML labels

inhibits · inhibitor23k+
inhibits · non-inhibitor39k+
metabolized_by (CYP substrate)1.9k+

Protein–protein2M+

physically_interacts_with1.5M+
interacts_with530k+

Gene & protein function, expression10M+

has_domain (protein→domain)2.6M+
participates_in (process + pathway)2.4M+
expressed_in2.2M+
enables (molecular function)1.2M+
located_in (cellular component)960k+
coexpressed_with450k+
part_of (pathway hierarchy)23k+

Gene regulation & wiring5M+

regulates (TF→gene)5M+
encodes (gene→protein)18k+

Metabolism (biochemical)290k+

reaction ↔ metabolite260k+
catalyzes (enzyme→reaction)36k+

Orthology & gene families410k+

orthologous_to (cross-species)250k+
paralogous_to (human)120k+
member_of (gene/protein→family)37k+

Drug effects & interactions2.2M+

drug_drug_interaction1.3M+
affects (chemical→gene)850k+
has_side_effect64k+

Clinical2M+

investigates_condition1M+
in_clinical_trial_for540k+
tests_intervention400k+

Retrosynthesis1.5M+

has_reactant600k+
derives_from600k+
has_product340k+

Genetics & gene–disease3.1M+

gene_associated_with_condition1.1M+
is_sequence_variant_of (variant→gene)1M+
variant_associated_with_disease640k+
associated_with_trait (GWAS)260k+
affects_expression_of (eQTL)28k+

Organism & taxonomy47M+

in_taxon (entity→organism)46M+
subclass_of (taxonomy tree)1.2M+

RNA & ncRNA330k+

rna_interacts_with (lncRNA→gene)130k+
mirna_targets (miRNA→gene)100k+
rna_binds_protein (RNA→protein)77k+
transcribed_from (transcript→gene)89k+
rna_compound_interaction (RNA→drug)1.4k+

Ontology & disease associations510k+

subclass_of (ontology)180k+
has_phenotype150k+
chem_associated_with_disease69k+
treats51k+
indicated_for (phase-scored)31k+
contraindicated_for30k+
exact_match647

Nothing untraceable.

Every edge carries source, evidence_type (experimental · computational · predicted · text_mined · curated), a native score, a negated flag, and evidence_count. Per-edge provenance and affinity measurements are stored separately, so both can be queried without denormalizing the core graph. Where sources disagree, the conflicting edges are all retained with their own evidence; positive evidence takes precedence only where a query has to resolve to a single sign.

Canonicalized and feature-complete: 39k+ duplicate nodes were merged into a single structured canonical form and their edges repointed, so no two nodes represent the same thing. Proteins carry amino-acid sequences for 98% of entries, and every structured chemical carries a machine-readable structure.

See for yourself

Real edges. Not illustrations.

A live edge pulled from the graph for every predicate below, with entity types, evidence, and sign: actual rows, not mockups.

PredicateExample edge (subject → object)EvidenceSign
bindsMequitazinechemical→HRH1proteinExperimental+
physically_interacts_withPARP1protein→MAPK7proteinCurated−
interacts_withHNRNPA2B1protein→SUMO2proteinComputational+
expressed_inFCGRTgene→material anatomical entityanatomyCurated+
participates_inPARP1protein→D4-GDI Signaling PathwaypathwayCurated+
located_inARHGEF7gene→cell cortexcell_componentCurated+
enablesGALNT7gene→carbohydrate bindingmol_functionCurated+
drug_drug_interactionProtriptylinechemical→CisatracuriumchemicalCurated+
affectsResiniferatoxinchemical→TRPV1geneCurated+
has_side_effectIsofluranechemical→JaundicephenotypeCurated+
investigates_conditionJump Start Plus COVID-19 Supporttrial→mental health wellnessdiseaseCurated+
in_clinical_trial_forCS1chemical→pulmonary arterial hypertensiondiseaseCurated+
tests_interventionSBRT or TACE for Advanced HCCtrial→DEBchemicalCurated+
has_reactantChloro N-alkylationreaction→Methyl alcoholchemicalCurated+
has_productSulfanyl to sulfinylreaction→CHEMBL4793427chemicalCurated+
derives_fromCHEMBL1885180chemical→CHEMBL1305819chemicalCurated+
has_phenotypediffuse palmoplantar keratodermadisease→Palmoplantar keratodermaphenotypeCurated+
gene_associated_with_conditionHSBP1L1gene→Thiel-Behnke corneal dystrophydiseaseComputational (60)+
variant_associated_with_diseaseNM_000355.4(TCN2):c.581-17variant→transcobalamin II deficiencydiseaseCurated+
is_sequence_variant_ofNM_000355.4(TCN2):c.581-17variant→TCN2geneCurated+
exact_matchHurler syndromedisease→mucopolysaccharidosis type 1diseaseCurated+
regulatesPOU2F1gene (TF)→TMEM9BgeneExperimental (both-confirmed)+
encodesGAD1gene→GAD1proteinCurated+
chem_associated_with_diseasemefloquinechemical→hearing loss, suddendiseaseCurated+
treatsRutinchemical→bacterial infectionsdiseaseCurated+
contraindicated_forPheniraminechemical→benign prostatic hyperplasiadiseaseCurated+
subclass_ofcollagen metabolic processbio_process→collagen biosynthetic processbio_processCurated+

Ready to put it to work?

The graph is live behind every run the co-scientist makes. Come see what it can trace for you.

Talk to us →Back to the platform