
Biochemistry Professor Peter S. Kim co-led research on a new machine learning-based approach to designing antibody drugs that was shown to improve on an FDA-approved SARS-CoV-2 antibody that had been discontinued due to its ineffectiveness. Photo by Steve Fish.
Proteins have evolved to be better at everything from contracting muscles to digesting food to recognizing viruses. To design better proteins, including antibodies, scientists repeatedly mutate amino acids – the units that proteins are arranged in – at different positions until the resulting protein functions better, such as triggering a stronger immune response or capturing carbon dioxide from the atmosphere more efficiently.
But the number of possible amino acid sequences is greater than the number of grains of sand on Earth, and finding the best proteins — and therefore the best potential drugs — is often expensive or impossible.
Scientists at Stanford University have developed a new machine learning-based method to more quickly and accurately predict the molecular changes that will lead to better antibody drugs.
Published in Science The approach, announced July 4, combines 3D structures of protein backbones with large-scale language models based on amino acid sequences to enable researchers to discover rare and desirable mutations in minutes that would normally be uncovered only through exhaustive experimentation.
The research team, led by Peter S. Kim, professor of biochemistry and laboratory researcher at the Sarafan Institute of Chemical Science, and Brian Hee, assistant professor of chemical engineering, showed they could improve a previously FDA-approved SARS-CoV-2 antibody that was discontinued in November 2022 due to its ineffectiveness against the new strain. Their approach improved its effectiveness against the virus by 25-fold.
“Much of the effort in AI and drug development has focused on amassing large amounts of data about how well a particular molecule performs a particular task, so that computers can learn well enough to design better versions,” Kim said. “What's remarkable is that we've shown that computers can still learn when we use that structure as a proxy for large amounts of data.”
“Now, more antibodies have the chance to really be optimized,” said Hie, who is also an innovation researcher at Ark Research.
Bend into shape
Faced with the challenge of finding the best amino acid sequence, scientists often make millions of them and then test them in miniature, simplified versions of living systems. The scientists hope that the best drug in the dish will also be the best drug for humans.
“It's a process of guessing and testing,” High said. “The goal of many intelligent algorithms is to take the guesswork out of it.”
To speed up this process, scientists have developed machine learning algorithms like ChatGPT, which are trained on millions of proteins’ amino acid sequences to predict desirable mutations.
But these models often point scientists to sequences that become unstable after being generated in the lab, or end up worse off than when they started.
This is partly because a protein's function depends not only on its sequence of amino acids but also on the 3D structure of that sequence: to trigger an immune response, for example, an antibody must be in the right shape to bind to a molecule on the surface of the virus.
The researchers believed that structure was the key to developing better prediction algorithms, so they narrowed down the long list of potentially beneficial mutations determined by their sequence-based large-scale language model to only those that preserve the 3D shape of the starting protein.
Testing Center
In December 2022, the research team tested this with a recently discontinued SARS-CoV-2 antibody treatment.
“The prevailing theory was that if you tried to improve on these antibodies, they would fail,” says Varun Shankar, a medical student and biophysics graduate student and lead author of the study. “The virus was very clever; as it spread through millions of people, it evolved to the point where it knew exactly how to mutate to evade these antibodies.”
Optimizing the protein using a purely sequence-based model only increased efficacy by a factor of two, but with a structure-based approach the team saw a 25-fold increase.
“We have finally caught up with the virus,” said Shankar, who is also a fellow in the Chemistry-Biology Interface Training Program at the Sarafan Institute of Chemical Engineering.
Teaching an old model new tricks
Most efforts to use AI to develop better medicines rely on “training” or “supervising” models, which involves generating vast amounts of data about the function and performance of unique protein sequences. This approach is time-consuming and results in models that are tuned to specific proteins to perform specific tasks.
The model does not require any input about the function of the protein or how well it functions, or even laboratory experiments: structure is so tightly coupled to function that the protein's coordinates become a proxy for performance.
In their study of COVID antibodies, they constrained not only the structure of the antibody itself, but also the structure of the antibody when it bound to the virus. From there, their model “learned” some of the rules of antibody binding without being taught them.
Early experiments show that the approach can be generalized to other kinds of proteins, such as enzymes that catalyze chemical reactions in the body. So far, the researchers have found that the model has shown scientists dozens of proteins, and on average, half of them are better than their starting point.
This tool could help us respond faster to emerging and evolving diseases, while also lowering the barriers to developing more effective medicines.
A more potent drug would require a smaller dose, meaning more patients could benefit from a given amount — a potentially breakthrough for diseases like HIV, where research has shown that high doses of antibodies given infrequently can protect patients from infection.
The team is making their models and code freely available to anyone.
“This is a great example of the power of deep learning to democratize the process of building better proteins,” Shankar says, “not only enabling new drug development but also opening up new areas of scientific exploration that were previously inaccessible.”
For more information:
Varun R. Shanker et al. “Unsupervised evolution of protein and antibody complexes using structure-based language models” Science (2024). DOI: 10.1126/science.adk8946
Courtesy of Stanford University
Quote: AI approach optimizes antibody drug development (July 5, 2024) Retrieved July 5, 2024 from https://phys.org/news/2024-07-ai-approach-optimizes-antibody-drugs.html
This document is subject to copyright. It may not be reproduced without written permission, except for fair dealing for the purposes of personal study or research. The content is provided for informational purposes only.
