“This is kind of a paradigm-shifting approach to accelerating discovery.” New machine learning models are being used to predict how molecules affect gene expression and select promising drug candidates for two difficult-to-treat diseases.
In a multi-institutional collaboration led by Michigan State University (Michigan, USA), scientists have created a machine learning-based drug discovery platform that can be used to screen large compound libraries and optimize lead molecules based on transcriptomic features. The potential therapeutics identified in this study were then tested in human cell lines and animal models, ultimately yielding promising new drug candidates for two difficult-to-treat diseases. They are hepatocellular carcinoma (HCC), the third leading cause of cancer-related death worldwide, and idiopathic pulmonary fibrosis (IPF), a rare chronic lung disease.
Identification of drugs that reverse the expression of disease-associated transcriptomic signatures has been extensively studied to identify drug repurposing candidates, but the potential remains de novo Drug discovery remains underdeveloped.
Implementing such approaches for screening very large compound libraries requires gene expression profiles of the compounds. These can be used to train machine learning models that can infer gene expression based solely on chemical structure. Despite recent successes demonstrating the potential of using this method in preclinical drug discovery, studies to date have only included commonly studied compounds and have not yet investigated novel compounds or lead optimization, an essential step in early drug discovery.
Integrating computational and experimental techniques to decipher neuronal heterogeneity
Here, Andreas Pfenning (Carnegie Mellon University, Pennsylvania, USA) presents the experimental and computational techniques he uses to investigate cellular heterogeneity in the brain.
In an attempt to fill this gap in the literature, the researchers introduce the ‘Chemical Structure-Related Gene Expression Profile Predictor’, or GPS. This is a drug discovery system for screening large compound libraries and designing new compounds that can override transcriptional phenotypes.
First, they trained GPS on millions of published experimental measurements covering more than 70 cell lines. This included gene expression change readouts for 978 landmark genes in four commonly studied cell lines the team focused on: MCF7, HEPG2, PC3, and VCAP.
We then used this model to screen a large pool of compounds to identify and validate promising candidates for multiple diseases, with a focus on HCC and IPF, where new effective treatments are urgently needed.
Researchers used human HCC cell lines Hep3B, HepG2, and Huh7, as well as HCC and IPF animal models, and IPF human lung tissue samples to identify and validate therapeutic candidates that reverse disease-related gene expression. We discovered two unique compounds for HCC and identified one repurposing candidate and one novel antifibrotic molecule for IPF.
By doing this, they demonstrated the potential to apply transcriptomics-based approaches as a means of discovering new therapeutic targets to treat diseases. In hopes of furthering future drug discovery efforts, the team developed a web portal to share their code and allow other researchers to use GPS for virtual compound screening.
“This is kind of a paradigm-shifting approach for people to drive discovery,” declared Bing Chen, one of the study’s senior authors. “We want more people to try this approach. But most importantly, we want people to actually use this approach to discover new treatments.”
“I think we have already proven that this platform can be applied to two completely different diseases,” added Xiaopeng Li, another senior author. “This means this platform can be used for other diseases as well, unlocking its potential.”
