Loading…
Loading…
Genomics & NGS · R&D demonstration, public data · Aug 2026
Maps five variant-effect tools onto ACMG/AMP PP3/BP4 evidence tiers, independently calibrated and held-out validated for three foundation models.
Genomic foundation models (AlphaGenome, Evo 2) and alignment-likelihood models (GPN-MSA) are usually benchmarked against classical variant-effect predictors using bare correlation or AUROC leaderboards. That doesn't tell a clinical geneticist whether a score is strong enough to use as PP3/BP4 evidence under the ACMG/AMP variant-classification framework. AlphaMissense and CADD already have published, ClinGen-vetted evidence thresholds (Bergquist et al. 2025, Pejaver et al. 2022); the three foundation models don't.
Curated ~10,000 ClinVar SNVs (≥1-star review status, temporal split: calibration snapshot clinvar_20250715, held-out snapshot clinvar_20260720), split into missense and non-coding tracks using ClinVar's own molecular-consequence field. Scored all five tools. Applied published ClinGen thresholds directly to AlphaMissense and CADD. For the three foundation models, independently derived thresholds using Pejaver et al. 2022's local posterior-probability calibration method, ported from their MATLAB reference implementation, using a gene-region-restricted gnomAD v4.1 reference sample to estimate each track's disease-variant prior. Every derived threshold was validated on the held-out set only — never reusing calibration data.
This is a reproduction of Pejaver et al.'s calibration methodology on our own curated data, not new ClinGen guidance — not benchmarked against a client baseline yet, and the derived thresholds should never be presented as official ACMG/AMP recommendations. PP3/BP4 was formally validated around missense variant prediction; applying it to non-coding/regulatory/splice variants, as the entire non-coding track here does, is closer to the research frontier than settled clinical practice. CADD's non-coding thresholds are extrapolated from a missense-only ClinGen calibration (no published calibration of CADD on non-coding variants exists) — its strong non-coding AUROC is a real, honestly-reported finding, but the evidence-tier percentages it's scored against are out-of-scope thresholds, not an independently validated non-coding calibration. AlphaGenome's score is a median(|raw_score|) aggregation over gene-restricted RNA-seq/splicing tracks, not a hand-picked gene+tissue combination the way DeepMind's own published examples use, which likely understates its real ceiling. Two known simplifications versus Pejaver et al.'s exact method: no gene-panel restriction (they required ≥1 pathogenic variant per gene) and no explicit tool-training-set exclusion.