Stanford expert defends AlphaGenome variant scores in detailed thread, rebutting "AI slop" criticism
On September 18, Stanford computational genomics professor Anshul Kundaje posted a long thread systematically responding to Steven Salzberg's criticism that DeepMind's AlphaGenome Atlas is "AI slop sweeping genomics," while giving a technical breakdown of the team's newly launched AVI (AlphaGenome Variant Impact) score. Overall conclusion: AVI is a competitive and interpretable variant prioritization tool, but it is neither a pathogenicity probability nor a first-of-its-kind paradigm; Google's marketing did overstate things, and the field consensus is that the regulatory sequence code of genes is far from "solved."
Confirmed
- Background of the dispute: Salzberg criticized AlphaGenome Atlas as "AI slop"; Kundaje posted a long thread rebutting this, citing Nature coverage. In closing, he said he looked forward to more substantive criticism from Salzberg, calling the dismissal of AG's variant effect scores as slop "pure nonsense."
- What AVI is: a genome-wide variant impact prioritization score that uniformly ranks all 9 billion possible variants with PHRED scaling—AVI 10 corresponds to the top 10%, 20 to the top 1%, 30 to the top 0.1%, and 40 to the top 0.01%. Kundaje specifically clarified that it is not a calibrated probability of whether a variant is pathogenic and should not be interpreted as a probability.
- Technical approach: at its core AVI is a classifier that treats common variants with frequency ≥0.1% as (noisy) neutral/benign proxies and rarer variants as potentially deleterious-enriched for training. Inputs are 18 features per variant: 10 AlphaGenome features (compressed from context-specific effect score vectors across thousands of tracks, covering chromatin accessibility, transcription factors, transcription, splicing, polyadenylation, 3D contacts, etc.) + the AlphaMissense protein-coding variant score + 2 conservation scores + 3 protein loss-of-function (LoF) annotations + an indel indicator, integrated via a linear hypernetwork into a single impact score.
- Evaluation findings: the AVI paper formally compared against other SOTA variant prioritization scores on multiple accepted benchmarks for rare and common variants, showing overall competitive performance—slightly ahead on some metrics and clearly better on others. Kundaje also noted the evaluations have biases.
- Interpretability: AVI can offer functional hypotheses for why a variant is prioritized (e.g., binding, accessibility), and researchers can drill down directly on the AlphaGenome model with interpretation methods to pinpoint the sequence features a variant disrupts (e.g., motifs) and the affected cell types and readouts.
- Prior art: learning variant impact scores using negative selection signals as proxies is not a new paradigm—the most direct comparison is the widely used classic scoring method CADD.
- Original AlphaGenome paper benchmarks: generally better than existing SOTA, but not a huge leap on many specific tasks; evaluations mostly covered perturbation experiments, synthetic libraries, and silver-standard causal variant sets.
Unconfirmed
- Kundaje acknowledged that variant prioritization genuinely needs more experimental validation and direct testing, but argued the level of evaluation and validation presented in the paper is sufficient to support its core conclusions.
Why it matters
- Kundaje pointed out that Sundar Pichai's claim that "we mapped the effects of 9 billion variants" is overstated: what was actually done was predicting effects across thousands of cellular contexts, not experimentally "mapping" them.
- He emphasized the paper authors' and the field's consensus: AlphaGenome and all similar models remain very far from "solving" the regulatory sequence code of genes; the paper is candid about its improvements and limitations, and while DeepMind's marketing language lacks nuance, it stays within reasonable industry norms. This dispute shows how model releases from high-profile AI companies get scrutinized in detail by the professional community for the balance between "marketing hype" and "genuine scientific value."
2026-09-18 ~ 2026-09-18 · 20 related posts
Primary sources
- Computational genomics expert rebuts 'AI slop' criticism of DeepMind's AlphaGenome — anshulkundaje ·
- DeepMind's AVI Score Uses Common-vs-Rare Variant Frequency as a Noisy Neutrality Proxy — anshulkundaje ·
- Stanford professor fires back: calling genomic variant effect scores 'AI slop' is nonsense — anshulkundaje ·
- [source] Computational genomics expert rebuts 'AI slop' criticism of DeepMind's AlphaGenome — anshulkundaje · 2026-09-18
- AlphaGenome Benchmarks: Better Than Prior SOTA Overall, But Not a Giant Leap in Most Cases — anshulkundaje · 2026-09-18
- AlphaGenome Paper Is Candid About Limits; the Marketing Hype Is Still Within Bounds — anshulkundaje · 2026-09-18
- Top Researcher: AlphaGenome Is Very Far From Solving the Gene Regulation Code — anshulkundaje · 2026-09-18
- What the New AVI Score Actually Is: A Genome-Wide Ranking, Not a Calibrated Pathogenicity Probability — anshulkundaje · 2026-09-18
- How AlphaGenome Compresses Thousands of Variant Effect Tracks Into 10 Features — anshulkundaje · 2026-09-18
- AlphaGenome Reduces Thousands of Variant Effect Tracks to 10 Per-Variant Features — anshulkundaje · 2026-09-18
- AVI Scores Each Variant Using 18 Features: AG, AlphaMissense, Conservation, LoF — anshulkundaje · 2026-09-18
- AVI Architecture: 18 Features Per Variant Fed Into a Linear Hypernetwork Ensemble — anshulkundaje · 2026-09-18
- [source] DeepMind's AVI Score Uses Common-vs-Rare Variant Frequency as a Noisy Neutrality Proxy — anshulkundaje · 2026-09-18
- AVI Follows an Established Paradigm, With CADD as Its Most Direct Comparison — anshulkundaje · 2026-09-18
- AVI Scores Are PHRED-Scaled and Traceable to Contributing Features — anshulkundaje · 2026-09-18
- Interpretability Methods Run Directly on AlphaGenome to Pinpoint Disrupted Motifs — anshulkundaje · 2026-09-18
- AVI Score Offers Interpretable Functional Hypotheses for Variant Prioritization — anshulkundaje · 2026-09-18
- AVI Is Competitive With SOTA on Established Variant Benchmarks — anshulkundaje · 2026-09-18
- AVI Variant Impact Score Matches SOTA on Benchmarks, Though Evals Are Biased — anshulkundaje · 2026-09-18
- Genomics researcher defends AI variant prioritization paper: validation already sufficient — anshulkundaje · 2026-09-18
- Stanford prof: Google overstated AlphaGenome claims, but calling it 'AI slop' is nonsense — anshulkundaje · 2026-09-18
- [source] Stanford professor fires back: calling genomic variant effect scores 'AI slop' is nonsense — anshulkundaje · 2026-09-18
1 near-duplicate retellings: anshulkundaje