Recent advancements in the field of molecular biology and biophysics have increased the importance of Artificial Intelligence (AI)-based approaches in order to deciphering evolutionary background of protein evolution. In order to complete their cellular tasks, proteins usually go through a folding (folding into a 3-dimensional structure) process. This process is strongly limited by natural selection, however, it is not usually enough for a protein to fold into a stable structure in order to fulfill complex processes such as delivering molecular signaling or enzymatic activity. Although the structural features of the protein universe have been largely mapped in the current scientific literature, the nature of functional constraints, which cannot be fully explained by physical folding rules and are referred to as dark matter or dark energy, is still not entirely understood. In this context, the question of how a balance is struck between stability and functionality during the evolution of protein sequences, and how this balance can be quantitatively expressed, remains important. While traditional physical models successfully explain the requirements for proteins to fold into a stable structure, they are sometimes inadequate in predicting the complex evolutionary pressures arising from biological functionality. In this context, AI technologies such as Large Language Models (LLMs) offer a more comprehensive picture of evolutionary processes by using statistical patterns learned from databases of protein sequences.
Galpern and his colleagues aimed to compare the effects of changes in protein sequences on folding energy and their effects on natural selection. The researchers aimed to develop an AI-supported computational framework to quantitatively determine the constraints that drive protein evolution. In the study, they aimed to predict the evolutionary suitability of changes in amino acid sequences using protein language models such as ESM-2 (Evolutionary Scale Modeling-2). By comparing the evolutionary energy (Ψevo) scores calculated through these models with physical folding energies, it was aimed to distinguish how proteins evolved not only for structural stability but also for functional requirements (catalysis, binding, etc.). This study revealed how Al models were used to predict the probabilities of variations in protein sequences and to map dark energy regions that could not be fully explained by physical rules.
In the study, a statistical mechanics-based computational and analytical approach was followed to analyze the effects of single-point mutations in protein sequences. In this context, the physical folding energies of the proteins (Efold) were determined using experimental deep mutational scanning data and modeling such as AWSEM (physically based coarse particle force field). In contrast, evolutionary energy scores (Ψevo) were calculated through protein language models such as ESM-2. The ESM-2 protein language model, featuring 650 million parameters and 33 layers, was trained on the UR50D dataset to predict protein evolution. Dark energy (Edark) is derived by scaling the difference between the physical folding energy cavity and the evolutionary energy gap of a protein with a suitable selection temperature. The analyses were conducted on large datasets involving enzymes and protein-protein interactions, and the effects of the variants were comparatively examined. The method used the masked marginal likelihoods of the model. In this process, a specific position in the protein sequence was artificially concealed (masked) and the model was asked to predict what would be the most suitable amino acid for that position in an evolutionary context. Using the log-probability ratios (logits) produced by the model, the evolutionary probability difference between mutant and wild-type amino acids was calculated and this value was defined as the evolutionary score (Ψevo). These obtained artificial intelligence-based scores were analyzed by comparing them with physical force fields such as AWSEM or folding energies (Efold) calculated with experimental data.
In conclusion, it has been determined that dark energy is largely concentrated in the functional regions of proteins, especially in active regions and binding interfaces. As a result of analysis with Al models, it was shown that protein language models could predict evolutionary constraints with high accuracy and detect functional regions that physical models could not predict. It was determined that the evolutionary scores obtained with the ESM-2 model strongly coincided with the points (dark energy) where evolutionary scores deviate from the physical folding energies, the active regions of the proteins and the binding interfaces. It has been concluded that these Al-based estimates are highly correlated with experimental data (e.g. Barstar-Barnase interaction) and can be used as a quantitative criterion of functional pressures in protein evolution.
Reference: Galpern, E. A., Bueno, C., Sánchez, I. E., Wolynes, P. G., & Ferreiro, D. U. (2026). Probing the dark energy in the functional protein universe. Proceedings of the National Academy of Sciences, 123(4). https://doi.org/10.1073/pnas.2531111123
Author: Dilara Susmuş
Editor: Elinsu Ak
-Bioinfocodes Scientific News Service-
News articles prepared by our team members, reviewing and compiling scientific research
published in journals with and impact factor greater than 20 (click here for the list).
