home

Evo2 Variant Effect Database
Home Worked examples About & fields Data sources & research Model comparison Updates Browse table SQL
Independent modules Variant Analysis Region/Region Set Analysis Trait Analysis
Guided exploration Guided Workflow

Zero-shot variant-effect scores for 6.48 million common human variants, from the Evo2 DNA foundation model.

GRCh38 · chr1–22 · website v1.2 · 2026-09-24

Release notes and update history

6.48 million human variants, scored by a DNA foundation model.

Look up a variant and see how extreme its score is for a variant of its kind, which biobank traits have stored associations, and what conservation and allele age say. Use three independent analysis modules, or follow a single study through the Guided Workflow. No SQL required.

How one Δ score is produced — and how to read it

1 · input 2 · model 3 · score REF window 2 kb total · about ±1 kb …GCTGGCACTGGGCG T GGCCGCA… ALT window alternate allele inserted …GCTGGCACTGGGCG C GGCCGCA… Evo2 7B · 40B autoregressive DNA LM REF score = −5,704.97 ALT score = −5,689.02 Δ = ALT score − REF score = +15.95 rs429358 · 40B avgRC · rounded scores 6 Δ per variant: 2 model sizes × 3 strand strategies 4 · reading the sign 0 Δ ≪ 0 ALT more surprising Δ ≫ 0 ALT less surprising Near zero: a limited model likelihood shift rs429358 · 40B avgRC · Δ = +15.95 5 · reading the magnitude Ranking uses |Δ| — how much the model's view of the sequence changes between the two alleles. |Δ| carries no direction. The sign of Δ does — and it is not a direction of phenotypic effect.
The 2-kb reference context is centred on the variant (approximately 1 kb on each side). The alternate sequence contains the alternate allele, which may be a substitution, insertion or deletion; the illustrated change is a substitution. The values shown are the stored 40B avgRC scores for rs429358. The sign diagram is schematic. The variant report gives model-specific |Δ| percentiles for context.

What the score can support

  • How constrained this base is under species-scale sequence statistics.
  • Which variants in a locus to prioritise for functional follow-up.
  • Whether a set of regions is more constrained than the genome as a whole.
  • How signed Δ or |Δ| relates to per-SNP heritability in the stored S-LDSC analysis.

What it may not support on its own

  • Not a pathogenicity call and not a calibrated probability.
  • No molecular mechanism — no gene, tissue, or direction of expression.
  • The sign of Δ is not the direction of phenotypic effect.
  • Not causality: inside an LD block a tag variant looks the same.

Three independent analysis modules

Start with the object you want to inspect. Each module works directly, without a disease selection or workflow session.

Independent module Variant Analysis → Six model scores and percentiles, core annotations, and available association or experimental records. Evidence stays grouped by its source and original study. Independent module Region/Region Set Analysis → Inspect a gene or genomic window, a predefined region set, or your own BED. Compare scores, conservation, allele age, frequency and functional strata. Independent module Trait Analysis → Browse existing single-annotation results by biobank and original phenotype: coefficient, 95% CI, z-score, heritability and within-biobank rank. Select signed Δ or |Δ|.

Region views: gene / single region · region set / BED · GWAS association tracks. Track overlap is descriptive and does not establish statistical colocalization.

Guided Workflow

Explore one explicitly selected GWAS study from its stored associations to nearby genes, a region set and individual variants. The independent modules remain available throughout.

One GWAS study phenotype · source · version P ≤ 5×10⁻⁸ or P ≤ 10⁻⁵ → Nearby genes TSS within ±1 Mb of a seed retain each seed–gene link → Region set gene body ±500 kb union · all scored members → Gene → Variant individual gene window exact allele · source evidence Optional: view Trait Analysis results where a reliable study mapping is available. No online S-LDSC rerun.
Overlapping gene windows are merged for set analysis. Selecting a gene opens its own window. The independent APOE example retains its original gene-query behavior.
Follow a trait from association to variant

Choose one study, inspect its gene windows, and open the region set for analysis.

Open Guided Workflow →
Variant report

Score in context, GWAS evidence, conservation, links out.

Variants in a gene

Auto-plots the gene's variants.

Plot a genomic region

GRCh38 coordinates → an interactive Δ landscape.

Search everything

Matches rsID, gene symbol, or amino-acid change.

Which score should I use?

Compare the six configurations across functional evidence: MPRA, QTL, GWAS fine-mapping and clinical annotations, with a separate S-LDSC sensitivity analysis. The frozen 14 September 2026 figures include their full captions and exact result tables; they describe the original analysis samples and do not establish a universal best configuration.

The default is Evo2_40B_AvgRC_Delta, the configuration used for the manuscript's main case studies. All six configurations remain available: two model sizes (7B and 40B) with forward-only (noRC), averaged reverse-complement (avgRC), or weighted reverse-complement (wtRC) scoring.

Performance varies across tasks, outcomes and analysis samples. The models differ in scale and training-data exposure, while the scoring schemes differ in how they combine positions and strands. Use the model-comparison figures to assess the relevant task, and report the configuration you use. Model scores are not interchangeable.

The score, six configurations and database fields →

Try an example

rs429358The APOE-ε4 variant: score in context plus every biobank trait it hits./variant APOEPlot a gene's common variants./viewer UKB hypertensive diseasesGCST90473520: inspect rs12509595 at FGF5, the manuscript's hypertension example, alongside 40B avgRC scores.FGF5 · rs12509595 FinnGen depression / dysthymiaF5_DEPRESSION_DYSTHYMIA: inspect rs2888295 at SNRK–ANO10, the manuscript's second association example.SNRK–ANO10 · rs2888295
6,475,578variants scored
2 × 3models (7B/40B) × strategies
4,772biobank GWAS linked
~99%noncoding (~37.9k coding)

~2.61M intergenic · ~2.18M intronic · ~1.12M ncRNA-intronic · gnomAD non-Finnish-European, MAF ≥ 5%, chr1–22 only.

Or browse by category

Functional region Coding effect Legacy functional/regulatory label Repeat class Open the full table →

Data & programmatic access

  • Browse the table — sort, filter, facet in the browser.
  • Run SQL — arbitrary read-only queries.
  • JSON API — the variant report as JSON; append .json to any table page.
  • About & feature scheme — the score, the six configurations, and database field definitions.

Research use only. Δ scores are uncalibrated zero-shot model outputs and are not validated for clinical or diagnostic interpretation. Co-occurrence of a strong score and a strong association identifies candidate functional variants, not causal ones.

License: Apache-2.0. Evo2 delta-likelihood scores over gnomAD NFE common variants (GRCh38, chr1–22).

Cite Evo2VED and the Evo 2 Nature paper · Update history

Browse · SQL · JSON API · About

Powered by Datasette · Data license: Apache License 2.0