Top
Best
New

Posted by utiiiD 19 hours ago

AlphaGenome Atlas: a high-resolution map of human DNA(blog.google)
https://deepmind.google/blog/alphagenome-atlas-a-predictive-...

https://deepmind.google.com/science/alphagenome/atlas

562 points | 121 comments
leopoldj 18 hours ago|
Videos. [2] is for the scientists to start using AlphaGenome Atlas from AntiGravity.

1. https://www.youtube.com/watch?v=U0aToL5C-bQ

2. https://www.youtube.com/watch?v=b2qw3rDNX0Q

adrian_b 15 hours ago||
In another HN thread about this AlphaGenome Atlas, someone has posted a link to:

https://www.science.org/content/blog-post/mutate-em-all-and-...

which comments the results of this study:

https://www.biorxiv.org/content/10.64898/2026.07.25.740675v1

That study has done in reality what the AlphaGenome Atlas does in fiction, but instead for a human they have done it for one of the simplest viruses.

So they have fuzzed the virus by mutating one by one each position of its DNA.

And various dedicated AI models all made poor predictions of the results of that experiment, which casts doubts about the value of the AlphaGenome predictive map.

A virus is much simpler than a human, but even for that simple virus the effects of most of the mutations could not be predicted. A half of the mutations had harmful effects, and for a half of those it is unknown for now why they were harmful.

For a human the uncertainty about the effects of a mutation will be far greater than for one of the simplest viruses.

derangedHorse 32 minutes ago||
> And various dedicated AI models all made poor predictions of the results of that experiment, which casts doubts about the value of the AlphaGenome predictive map

I would not group AlphaGenome into the pile of failed predictions of other models. AlphaGenome deserves to get evaluated based off its own merits.

John7878781 14 hours ago||
Yep. Sequence-to-function models are still very limited. AlphaGenome Atlas, despite the flashy branding, is unlikely to provide significant benefit to researchers.
tsoukase 5 hours ago||
May be that's the reason for the alpha naming. We are waiting for a stable release (just kidding).
DoctorOetker 16 hours ago||
Not a word about promoter sequences.

Imagine cellular activity as an industry zone, its not just what you can or can not make, its also 'for what concentrations of chemical species, what transcription rates should be used' so apart from the discrete Mendelian aspects (like what eye color or what have you) there is also a concensus sequence and deviations from consensus. They mention the dataset captures non-coding DNA, which should imply promoter sequences. Will it be possible to query the atlas for joint probabilities of promoter and putative target protein occurence in human genomes?

Personalized medicine could never credibly take off as long as promoter sequences were excised before sequencing!

dekhn 11 hours ago|
Eye color is not discrete Mendelian. That's only correct to the first order.

Also Mendelian has little to do with promotor sequences or differential transcription in deviations.

Stevvo 13 hours ago||
Don't be put off by the box asking for your "affiliation". I wrote "None", clicked submit and it took me straight to the Atlas.
r0ze-at-hn 7 minutes ago|
The agreement does pretty much state you can't use this for anything useful.

As someone who regularly investigates whole genomes I would love to use this as a tool on novel mutations. These folks are the edge cases no one else could figure out that I get a crack at. Beyond the DNA we have the symptoms and lab work and I can usually narrow it down to a handful of guesses, but it sure would be nice to use this to help rank where to invest efforts.

For now i'll treat it as just another fun Google project that might come out of beta one day (or not).

RobotToaster 17 hours ago||
Can this be used with a 23andMe genome to find pathogenic mutations?
MisterMunchkin 16 hours ago||
23andMe and similar companies don't transcribe your entire genome because that would cost way more than they charge you. They just sample a few tiny sections of it.
willturman 14 hours ago|||
A few as in tens of thousands.

23andMe used a custom Illumina Infinium microarray designed around segments of particular interest.

https://www.illumina.com/products/by-brand/infinium.html

iooi 14 hours ago|||
fwiw sequencing your entire genome only costs $400 or so from providers like sequencing.com
wraptile 7 hours ago||
Thank you for the recommendation! I've been on a lookout for an affordable full sequence and 400$ is very reasonable. Very excited to just have my DNA on my own computer - this feels like the good kind of sci-fi :)
vibrio 1 hour ago||
What are their data privacy and security policies? That is an important consideration.
vintermann 14 hours ago|||
Probably not any 23andMe haven't already told you about. They test a limited set of SNPs, balancing between ones thought useful for genealogy, ones useful for ethnicity estimates and ones thought useful for health-related things (the latter they would like to make their main selling point, the two former are really all commercial DNA services' bread and butter).

It's unlikely that they would luck into testing some unknown SNP which turned out to be relevant for disease.

MadrasTh0rn 15 hours ago|||
23andMe tests SNP's (single nucleotides) that are inferred to be significant in protein function/epigenitics.

Those SNP's i believe are testd from primers

so what 23andMe does is specifically on the back of previous research and afaik their data isnt technically clinically significant as most findings need confirmation or more tests.

dekhn 17 hours ago||
Not really, no.
shevy-java 16 hours ago||
Why not? There is no logical reason as to why this would not work, IF it works in the first place, which I don't know. In theory the problem space here is finite, so there is of course a way to predict everything. Whether this is the case right now - who knows; I probably don't think it is currently ready. But eventually it will be. And it should not be in the hands of private companies.
dekhn 16 hours ago|||
So it sounds like you're coming to this with very little knowledge about biology. I encourage reading up on modern challenges in pathogenic prediction, especially with regards to SNPs: on their own, with the exception of a few diseases, individual SNP predictions are meaningless in terms of actual pathogenicity.
chorizo 15 hours ago|||
Because 23andme does not sequence your genome, only substring matches linked to specific gene variants.
SubiculumCode 14 hours ago||
Alpha Fold has continued to impact the field of protein networks, but I do hear that not every one of the deep learning biology models from Google/Deep Mind and others have made equivalent impact or had as lasting relevance in their respective domains..some have performed more poorly than other available models. I'd love to learn more about this, but this has mostly come from little snips of conversations here and there, in person and online, but I haven't seen anything comprehensive in terms of evaluating their impacts overall
bonsai_spool 17 hours ago||
This may not be anything new but it makes using several Google/DeepMind resources a lot less painful.

I'm comfortable programming but others who also do mol bio may be less so or may not recognize when Claude is going off the rails.

John7878781 15 hours ago||
People are upvoting this because it has the “Alpha______” prefix. Meanwhile, everyone in the field of genomics knows that AlphaGenome provides essentially zero improvements over the previous SOTA, Borzoi…
atorodius 15 hours ago||
> a database that predicts the effects of every possible single nucleotide variant in the human genome. We used the AlphaGenome AI model to pre-calculate the regulatory impact of all 9 billion single-letter genetic changes, resulting in a massive, 1-petabyte dataset.

This is for a database, no? While Borzoi is a model?

> Here, we introduce Borzoi, a model that learns to predict cell-type-specific and tissue-specific RNA-seq coverage from DNA sequence.

https://www.nature.com/articles/s41588-024-02053-6

monocasa 13 hours ago||
Any model can be expressed as a database.
baq 12 hours ago||
Kolmogorov looks at this with a ‘duh’ face
ericmay 14 hours ago|||
> Meanwhile, everyone in the field of genomics knows that AlphaGenome provides essentially zero improvements over the previous SOTA, Borzoi…

Can you elaborate on this? I'm confused why Google would build something that provides zero improvements over SOTA, Borzoi... as you mention. I'm not familiar with this field, just curious.

dekhn 14 hours ago|||
Speaking as somebody who has worked within Google Research before: the researchers are under tremendous pressure to publish SOTA and sometimes they juice their results a bit to look competitive when they can't match. This is not uncommon in the field- it's remarkably easy to edit a paper to make yourself look good by omitting information.
maxall4 13 hours ago||
One of the most egregious cases of this, in my opinion, is only publishing metrics that cover part of the confusion matrix. “The false-negative rate? That could not possibly matter for a variant effect prediction model; why would we include that in the paper?” Example: AlphaMissense.
AISlopCannon 13 hours ago|||
The use of an exact quote in an ungrammatical fashion is a bit of a language model smell. I can’t help but be reminded of the purely nonsensical AI interview answers. “It’s a pleasure to meet you, Chick Bongo”
ericmay 13 hours ago||
Creating an account 12 minutes ago (from the time of this posting) to comment on how another comment seems to you like a "language model smell" is itself, a language model smell, or a scammer.

Please stop accusing or hinting at others being a language model or bot. Not only is it a dumb waste of time, it's wrong in this instance and you are not only going to continue to be wrong but you have no way to prove or demonstrate that any single post comes from a bot nor the ability to do anything about it if you did in fact believe some comment to be attributed to a bot.

mbeavitt 10 hours ago|||
They did benchmark the model and beat the SOTA on every metric though
scottLobster 14 hours ago|||
Is it so bad to have another entrant, especially with the resources Google could bring to bear?

Imagine if the Apple EV had actually happened, you think the EV enthusiasts would roll their eyes like you are?

weedfroglozenge 11 hours ago||
Are you really taking a holier than thou approach on a google article?
zmmmmm 11 hours ago||
Is this just Google precomputing Alpha genome values - which were already accessible via API and making them available as another API (presumably more broadly)? Or is there actually new information?
mbreese 11 hours ago|
That’s my reading. (That this is a cached database of Alpha values)
mchusma 16 hours ago|
This has Demis written all over it. There is a great video of him with AlphaFold chatting with the team about releasing some results, and he asked something like “what if we just do them all?”

Very excited to see that happen here.

John7878781 15 hours ago||
I don’t think Demis played a big role in this. It was mainly Ziga Avsec who developed Enformer (the first actually decent sequence-to-function model), and then AlphaGenome.
dekhn 16 hours ago|||
This has nothing to do with AlphaFold at all- not sure if you were implying that. (the scientific contriution is welcome, but it's not particularly significant)
shevy-java 16 hours ago||
Why are you excited about this?
More comments...