Interpreting the biological effects of genetic variation is regarded as one of the fundamental challenges of modern biology. More than 98% of the variations in the human genome reside within protein non-coding regions, and these variants can influence gene expression, chromatin architecture, and RNA processing (splicing) through intricate mechanisms. Existing computational approaches developed to predict these effects are generally forced to choose between two principal limitations: they can process long genomic sequences but provide low resolution, or they offer high (single-nucleotide) resolution while being restricted to analyzing only short DNA segments. This situation constrains the comprehensive deciphering of the genetic regulatory code and the precise modeling of long-range genomic interactions.
In this study conducted by Google DeepMind, the objective was to overcome the limitations of resolution and coverage encountered in predicting the effects of genomic variants. The researchers developed an integrated model named AlphaGenome, aiming both to capture long-range genomic interactions and to generate predictions at single-base (base-pair) precision. In contrast to existing approaches, this work sought to model diverse biological processes (modalities), such as gene expression, chromatin accessibility, and splicing, within a unified framework and simultaneously. In this way, the gap between the isolated successes of specialized models and the broader scope of general models was intended to be bridged, enabling more accurate predictions of the molecular consequences of regulatory variants.
The study employed AlphaGenome, a deep learning (DL)-based architecture. The model accepts DNA sequences of 1 megabase (1 million base pairs) as input and is trained on both human and mouse genomes. The method combines a U-Net-like architecture with Transformer blocks, designed to model both local sequence patterns and interactions between distant genomic regions. The model is configured to predict thousands of functional genomic tracks, including gene expression, transcription factor binding, histone modifications, chromatin contact maps, and RNA splicing events. During training, a distillation strategy was applied, in which the knowledge of multiple teacher models was transferred to a single student model in order to enhance performance and stability.
As a result, the AlphaGenome model was reported to achieve a substantial performance improvement over existing methods in predicting the effects of genetic variants. In 25 out of 26 variant effect prediction benchmarks, the model was observed to match or surpass the strongest external models. Notably, the mechanisms of clinically significant variants surrounding the TAL1 oncogene were reported to be accurately summarized by the model across all modalities. The findings indicate that the model is capable of integratively analyzing alterations in splicing, gene expression, and chromatin structure. These results are considered to offer researchers a novel perspective for interpreting the molecular-level functional consequences of genetic variants and for generating hypotheses to guide future experimental investigations.
Reference: Avsec, Ž., Latysheva, N., Cheng, J. et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature 649, 1206–1218 (2026). https://doi.org/10.1038/s41586-025-10014-0
