Deep learning approaches have transformed how scientists predict the activity and function of DNA sequences in the genome. A new AI tool called Corgi (Context-aware Regulatory Genomics Inference), developed by Martin Vingron's bioinformatics group at the Max Planck Institute for Molecular Genetics allows researchers to predict how the genome behaves across different cell types. The findings could help improve efforts to mimic the activity of cells virtually, with potential applications in basic research and medicine.
Subscribe to our newsletter for the latest sci-tech news updates.
Although different cells in the body may appear similar, each person has a unique genetic makeup because of small genetic variations. These variants can alter gene expression in cells. Researchers are increasingly interested in these variants because they may be involved in disease mechanisms and could inform personalized medical interventions. However, only a few variants are functionally relevant. In recent years, AI-based tools have been developed to predict the effects of these variants from sequence data and prioritize the most important ones.
Where earlier models fell short
Many of these tools have fundamental limitations. Some cannot extrapolate beyond their training data because they fail to consider the cellular context that determines how genes are switched on or off. "With Corgi, we have created a model that overcomes many of these limitations. It can correctly predict measurements related to genome activity, including gene expression, chromatin accessibility and histone modifications," says Ekin Deniz Aksu, first author of the study published in Nature Communications.
Earlier models typically disregarded this cellular context, relying solely on sequence. These models were trained using large data sets that linked DNA sequences to outputs, such as RNA activity or chromatin accessibility. However, DNA does not act alone but with a plethora of regulatory molecules. The same DNA sequence may behave differently in a neuron than in a liver cell, for example, because each cell contains a different mix of regulatory molecules. "To train our model, we created an architecture that more closely resembles actual biological reality in a cell by integrating expression data from a large set of regulatory genes," Aksu says.
From missing data to virtual cells
An improved version of the tool, called Corgi+, can infer epigenetic information from RNA sequencing data. With further adjustments, researchers could use the tool to fill major gaps in data sets without conducting every experiment directly, particularly with regard to rare cell types, embryos and limited patient tissue.
The scientists hope their approach will also improve virtual cell models. These systems aim to mimic the effects of gene perturbations at a cellular level, but they do not incorporate information about DNA activity. "We believe that our model could help merge sequence-to-function models with virtual cells in the future," Aksu says.
In the long run, such systems could help scientists screen for drug candidates virtually and understand how certain mutations alter cellular processes and cause complex disease phenotypes.
More information: Ekin Deniz Aksu et al, Context-aware sequence-to-function model of human gene regulation, Nature Communications (2026). DOI: 10.1038/s41467-026-75527-2
Provided by Max Planck Society
This story was originally published on Phys.org.