In the rapidly evolving field of medicine, new artificial intelligence models may be trying to change the way mRNA-based drugs and vaccines are designed. Developed through a collaboration between the University of Texas at Austin and pharmaceutical company Sanophy, the tool helps researchers predict how efficiently different mRNA sequences will produce proteins in the body. Its capabilities can significantly reduce trial and error in treatment design and speed up the development of life-saving treatments.
The scientists named this tool ribbon. Use artificial intelligence to predict translation efficiency. A method in which a cell can convert a strand of mRNA into a protein. The tool is based on a deep learning system derived from over 10,000 ribosome profiling experiments. These experiments spanned 140 different human and mouse cell types, generating 3,819 data sets, forming the most detailed translation efficiency atlas.
Cracks in the code of protein production
Cells produce proteins through processes that involve DNA, mRNA, and ribosomes. First, instructions for creating proteins are copied from DNA to messenger RNA. These mRNA strands then enter the ribosome, the protein factory of the cells. Here, instructions are used to assemble the amino acid chains into proteins. This process is not always easy, particularly for therapeutic purposes.
The sequence of mRNA can affect how well the ribosome affects its reading and translation. Until now, scientists have had limited tools to predict their efficiency. Many relied primarily on the characteristics of the 5' untranslated region (5' UTR) of mRNA. However, protein production is influenced by many sequencing functions, such as how codons are arranged and how ribosomes interact with their sequences as ribosomes migrate.
That's where the ribbon stands out. Unlike older models, we also consider how dinucleotide, trinucleotide, and codon positions throughout the sequence, not just the 5'UTR, but also the position of dinucleotides, trinucleotides, and codons affect protein production. This means that we can predict how the structure and arrangement of mRNA features affects the ability of cells to make proteins.
“We're a great leader in the project,” said Can Can Canik, an associate professor of molecular biological sciences at UT Austin and one of the project's leaders. “That is the value of curious research, which builds the foundation of progress, like ribbons.
Related Stories
From data to discovery
Before building the AI model, researchers at UT Austin and Sanofi collected data from public scientific experiments. In these experiments, we measured how cells efficiently convert different mRNA sequences into proteins in the body.
This work required attention to UT Austin's accuracy and the undergraduate researchers involved. They reviewed the experimental data and manually corrected missing or incorrect information. This cleaned and validated dataset, the dataset named Ribobase, formed the basis for training the Ribonn model.
The development of the model took several years of collaboration between academic researchers and industry researchers. Key contributors include Can Cenik and Vikram Agarwal, head of Sanofi's mRNA platform design data science. Other contributors include Logan Persyn, UT graduate students in computer science, and researchers Dinghai Zheng and Jun Wang, at Sanofi. The discovery of UTs affecting Office has helped unite academic and industry teams under a formal research agreement.
The technical aspects of the model are just as impressive as the biological aspects. The Ribbon is a multitasking deep convolutional neural network. This type of AI is often used in computer vision and natural language processing. Learn patterns from mRNA sequences. This model recognizes how small sequence features affect the overall process of protein translation. It also captures biological principles such as ribosome processability and tRNA abundance.
These factors affect how easily they match ribosome movement and amino acids. The team received support from the National Institutes of Health and the Welch Foundation. I also used a Lonestar6 supercomputer at UT's Texas Advanced Computing Center. Using these resources, they trained and tested ribbons on an unprecedented scale.
New tools for the next generation of medicine
In the trial, the ribbon surpassed the previous model by a wide margin. It often provides twice as accurate as predicting translation efficiency across many different cell types. This level of accuracy can revolutionize mRNA therapy. It opens the door to more targeted drug design, allowing scientists to predict not only how much protein cells will produce, but which cells will produce it.
“Maybe next-generation therapy was needed to make proteins in the liver, lungs or immune cells,” explained Cenik. “This opens up the opportunity to modify mRNA sequencing to increase the production of proteins in that cell type.”
This type of control may prove particularly useful in the treatment of cancer, infection, or genetic disorders. This is a condition in which targeting the right organization is important for success. Instead of relying solely on trial and error testing, researchers can use ribbons to pre-model potential treatments and identify the most effective options before entering the lab.
Additionally, ribbons can be used to study base-modified therapeutic RNAs that are commonly used in real-world therapy. These are specially designed versions of mRNAs that can resist degradation and help reduce the immune response.
Understanding how they behave within the cell allows scientists to fine-tune them to enhance their effectiveness. This model also provides insight into how evolutionary forces form mRNA sequences. It is possible to clarify why certain patterns of 5'UTR are conserved across species, indicating how translation efficiency led to natural selection.
Revealing shared biological language
The second paper is built on the same dataset and provides broader insights. mRNAs with similar biological functions tend to translate at similar levels regardless of cell type. For years, scientists have known that relevant genes are consistently transcribed into mRNA in a coordinated manner.
It is clear that this also modulates the process by which cells convert these mRNAs into proteins across different types. This reveals the general regulatory language that links mRNA production, stability, localization, and translation. By deciphering this language, scientists can design better treatments and understand how cells maintain internal balance and function.
This study will improve treatment design and provide a basic understanding of how cells function. “When I started this project over six years ago, there were no obvious applications,” recalls Cenik. Scientific curiosity has led to discoveries that have great significance for both science and medicine.
With tools such as Ribonn, personalized medicine has become less dependent on guesswork and more accurate predictions. Researchers can start with data-driven models to create better mRNA sequences and provide targeted therapy more quickly.
