2026-09-08

DeepMind Precomputes Every Possible Single-Letter DNA Mutation — and It Already Found a Real Disease Variant

AIScienceOpen Source🌍 Global

Google DeepMind announced AlphaGenome Atlas today: a precomputed database of predicted molecular effects for 9 billion single-nucleotide variants — every possible single-letter change in the human genome — built on top of AlphaGenome, the variant-effect prediction model DeepMind released earlier. The dataset is 1 petabyte, which DeepMind says is "more than 30 times larger than the AlphaFold Database," and it's free for academic use through a website portal, an API, and a skill inside Google Antigravity, DeepMind's agentic platform. Commercial access is coming "soon" via Google Cloud; the underlying AlphaGenome model itself is already available academically on GitHub and commercially through Cloud's Model Garden.

What the Atlas actually contains

Testing all 9 billion possible variants in a lab is not feasible, which is the whole premise here: DeepMind precomputed AlphaGenome's predictions across the genome instead of leaving researchers to query one variant at a time. Each variant gets thousands of molecular-effect predictions spanning gene regulation across hundreds of human and mouse cell types and tissues. On top of that, DeepMind introduces the AlphaGenome Variant Impact (AVI) score — a single number per variant that combines AlphaGenome's regulatory predictions with AlphaMissense, DeepMind's earlier model for scoring protein-altering variants. The AVI score is designed to work across both the roughly 2% of the genome that codes for proteins and the remaining 98% that doesn't but still regulates gene activity and carries most trait-associated variation. Each AVI score also comes with feature attributions — a breakdown of which specific biological processes (splicing, chromatin accessibility, conservation, and others) are driving that variant's predicted impact — plus a separate catalogue of more than 2,500 recurring DNA sequence motifs and where they occur genome-wide.

The strongest evidence here is a validated disease finding, not a benchmark score

The most concrete result in the announcement comes from a collaboration with the GREGoR Consortium. Researchers Laura Covill and Anne O'Donnell-Luria at the Broad Institute used the AVI score to re-prioritize variant candidates in an unsolved rare-disease case and found a variant in DNM1, a gene strongly linked to epileptic encephalopathy, that earlier analysis had overlooked. AlphaGenome's own predictions didn't just flag the variant — they specified the mechanism: it creates an incorrect splice site that produces an abnormally extended protein. That specific, falsifiable mechanistic prediction was then experimentally validated, and the same screen turned up nearby variants with similar effects. That's a materially stronger form of evidence than most model-launch case studies get to claim: not just a plausible-sounding output, but a named mechanism checked against a wet-lab result.

A population-scale result that's real but should be read as one team's finding, not a peer-reviewed benchmark

Separately, Gareth Hawkes, a Medical Research Council fellow at the University of Exeter, applied the Atlas to whole-genome data from more than 54,000 UK Biobank participants, grouping rare non-coding variants by their predicted molecular effects to cut through statistical noise that normally buries these signals. DeepMind's write-up reports this approach surfaced 22% more non-coding genetic associations than would otherwise have been detectable, including regulatory variants tied to circulating levels of PLA2G7 (linked to aging) and EGLN1 (a cellular oxygen sensor). A further pass restricted to the top 1% of Atlas-flagged non-coding variants identified 19 genetic regions potentially associated with body mass index. These are genuinely substantive, checkable numbers from an outside academic collaborator rather than DeepMind's own internal evaluation — but they're also reported here secondhand, inside DeepMind's own announcement, with no link yet to a published, peer-reviewed paper describing Hawkes' methodology in full. Worth treating as a strong early signal rather than a settled, independently reviewed result until that publication exists.

What's missing: no benchmark number behind "best-in-class"

DeepMind states that "the AVI score provides best-in-class performance across many variant pathogenicity and rare disease benchmarks," without naming which benchmarks, what the comparison models were, or what the actual scores were. That's a meaningfully weaker form of disclosure than the two case studies above, and it's the same kind of unqualified superlative this blog has flagged before in other Google DeepMind science releases — where a headline claim ("closed-loop execution," here "best-in-class") held up unevenly once the underlying data was checked (a length-adjustment needed to sustain one benchmark win there, an oxidized and structurally unconfirmed sample undercutting another). Nothing in this announcement is demonstrably wrong the way parts of that one were, but the absence of a single concrete number behind "best-in-class" is a real gap, not a stylistic quirk — it's the one claim in this release that isn't independently checkable from what's disclosed.

The AlphaFold Database precedent is worth taking seriously

DeepMind explicitly frames the Atlas as following the AlphaFold Database's playbook: that resource grew from roughly 190,000 experimentally determined protein structures to more than 200 million predicted ones in 2022, and became foundational infrastructure across structural biology largely because it was free, comprehensive, and usable by researchers without coding experience. If AlphaGenome Atlas follows a similar trajectory, the free non-commercial access now, with paid commercial access "coming soon" on Google Cloud, is the same two-track model DeepMind has used before: build the free version into a standard research tool first, monetize the version aimed at industry second. This blog's own tracking of AI-for-science capital has noted that money in this space concentrates heavily in drug discovery and materials; a free genome-wide variant atlas that becomes standard infrastructure for rare-disease research is a plausible way to build the customer base a commercial version later sells into.

What to expect next

  • Watch for the benchmark numbers behind "best-in-class." A specific comparison table would turn DeepMind's strongest unqualified claim into a checkable one; its absence here is worth tracking as future releases either fill the gap or repeat the same superlative.
  • Watch for peer-reviewed publication of the UK Biobank and rare-disease findings. Both are compelling as reported, but neither is yet independently reviewed in the form DeepMind's announcement presents it.
  • Watch the Google Cloud commercial pricing when it lands. The free academic tier is the AlphaFold Database playbook; what a paid tier costs, and who it's priced for, will say more about DeepMind's actual business model here than the free portal does.
  • Watch adoption outside the named academic partners. Broad Institute, Stowers Institute, and University of Exeter are strong initial validators; how quickly independent labs outside DeepMind's own collaborator network start publishing results using the Atlas is the real test of whether this becomes standard infrastructure the way the AlphaFold Database did.