Caroline Wooler is Andrew (1956) and Professor Elna Viterbi of Engineering at MIT. Professor of Electrical Engineering and Computer Science (IDSS) at the Institute of Data, Science and Social Research. Director of the Eric and Wendy Schmidt Center at MIT and Harvard's Broad Research Institute.
Wooler is interested in all the ways scientists can reveal the causal relationships of biological systems. In this interview, she discusses machine learning in biology, a ripe field of problem solving, and cutting-edge research that emerges from the Schmidt Center.
Q: The Eric and Wendy Schmidt Center have four different focal regions composed mainly of four natural-level biological tissues: proteins, cells, tissues, and organisms. What is the current machine learning landscape that has been the right time to tackle these specific problem classes?
A: Biology and medicine are currently undergoing a “data revolution.” The availability of large and diverse datasets, from genomics and multiomics to high-resolution imaging and electronic health records, makes this the right time. Inexpensive and accurate DNA sequencing is real, advanced molecular imaging is becoming routine, and single cell genomics allow for the profiling of millions of cells. These innovations, and the large datasets they generate, have brought us to the thresholds of a new era of biology. It can move beyond characterizing biology units (all proteins, genes, cell types, etc.). map.
At the same time, over the past decade, machine learning has made remarkable advances in demonstrating the advanced capabilities of text understanding and generation in models such as BERT, GPT-3, and CHATGPT, while multimodal models like Vision Transformers and Clip have achieved human-level performance in image-related tasks. These breakthroughs provide a powerful architectural blueprint and training strategy that can be adapted to biological data. For example, trans can model genome sequences similar to language, while vision models can analyze medical and microscopic images.
Importantly, biology is not just a beneficiary of machine learning, but also a key source of inspiration for new ML research. Just as agriculture and breeding have spurred modern statistics, biology could stimulate new, perhaps even deeper paths in ML research. Unlike areas such as recommendation systems and internet advertising, there are no natural laws to discover, prediction accuracy is the ultimate measure of value, physically interpretable in biology, and causal mechanisms are the ultimate goal. Furthermore, biology boasts genetic and chemical tools that allow perturbation screens at an unparalleled scale compared to other fields. These combined features allow biology to realize its own biology to benefit greatly from ML and serve as a deep well of its inspiration.
Q: Taking a slightly different tack, is the biology problem really resistant to the current toolset? Are there any areas of your illness or wellness challenges that you feel ripe for problem solving?
A: Machine learning has shown significant success in domain-wide prediction tasks such as image classification, natural language processing, and clinical risk modeling. However, in biological sciences, prediction accuracy is often insufficient. The basic questions in these areas are causal in nature. How do perturbations to specific genes or pathways affect downstream cellular processes? What mechanisms can intervention lead to phenotypic changes? Traditional machine learning models, primarily optimized to capture statistical associations of observational data, often fail to answer such intervention queries. Biology and medicine need to stimulate new fundamental developments in machine learning.
The field is equipped with high-throughput perturbation techniques such as pooled CRISPR screens, single-cell transcriptomics, and spatial profiling that generate rich datasets under systematic interventions. These data modalities naturally seek to develop models that support causal inference, active experimental design, and representational learning in settings with complex, structured latent variables. From a mathematical perspective, this requires addressing core problems regarding discriminability, sample efficiency, and integration of combinations, geometry, and stochastic tools. Addressing these challenges not only unlocks new insights into the mechanisms of cellular systems, but I think it will push the theoretical boundaries of machine learning.
Regarding basic models, the consensus in this field is that, like what ChatGPT represents in the linguistic domain of biology, a kind of digital organism that can simulate all biological phenomena, we are still far from creating holistic foundation models of biology across scales. New basic models appear almost every week, but these models have previously specialized in specific scales and questions, focusing on one or several modalities.
Great progress has been made in predicting protein structure from sequences. This success highlights the importance of iterative machine learning challenges, such as CASP (Critical Evaluation of Structural Prediction), which contributes to the benchmarking of cutting-edge algorithms for protein structure prediction and helps to drive improvement.
The Schmidt Center organizes challenges to raise awareness in the ML field and advance the development of ways to solve causal prediction problems that are of great importance to biomedical sciences. As the availability of single-gene perturbation data at the single-cell level increases, it is a solvent-enabled problem to predict the effects of single or combined perturbations and predict that perturbations can drive the desired phenotype. The Cell Perturbation Prediction Challenge (CPPC) aims to provide a means to objectively test and benchmark algorithms for predicting the effects of new perturbations.
Another area where the field has made significant advances is disease diagnosis and patient triage. Machine learning algorithms integrate different sources of different patient information (data modalities), generate missing modalities, and identify patterns that can be difficult to detect patients based on disease risk. We need to remain cautious about the potential bias in model prediction, the risk that models will learn shortcuts instead of true correlations, and the risk of automation bias in clinical decision-making, but I think this is an area where machine learning has already had a major impact.
Q: Let's talk about some of the headlines that have come up from Schmidt Center these days. Do you think people should be particularly excited?
A: In collaboration with Dr. Fei Chen of the Broad Institute, I recently developed a method to predict the intracellular location of an invisible protein called Pups. Many existing methods can only make predictions based on the specific protein and cell data they were trained to. However, puppies combine protein language models with image in-painting models to utilize both protein sequences and cell images. We demonstrate that protein sequence input allows for generalization to invisible proteins, and that cell image input captures single cell variability and allows for cell type-specific predictions. This model can learn how closely each amino acid residue is related to predicted subcellular localization and predict changes in localization due to mutations in protein sequences. Because protein function is strictly related to subcellular localization, our predictions may provide insight into the potential mechanisms of disease. In the future, we aim to extend this method to predict the localization of multiple proteins within cells and to understand protein-protein interactions.
Together with Professor GV Shivashankar, a longtime collaborator at EthZürich, we have shown that simple images of cells stained with fluorescent DNA intercalating dyes can generate a lot of information about the state and fate of healthy and diseased cells when combined with machine learning algorithms to label chromatin. Recently, we have demonstrated a deep link between chromatin tissue and gene regulation by developing Image2reg, a method that promotes this observation and allows for the prediction of invisible genetically or chemically perturbed genes from chromatin images. Image2REG utilizes a convolutional neural network to learn beneficial representations of chromatin images of convolutional cells. We also employ graph convolution networks to create gene embeddings that capture regulatory effects of genes based on cell type-specific transcriptome data and integrated protein-protein interaction data. Finally, we learn the map between the resulting physical and biochemical representations, allowing us to predict perturbed gene modules based on chromatin images.
Additionally, we have recently completed the development of a method to predict the outcome of invisible combined gene perturbations and identify the types of interactions occurring between perturbed genes. Morph can guide the most beneficial perturbation design for Lab-in-a-loop experiments. Furthermore, attention-based frameworks allow us to demonstrate that our methods identify causal relationships between genes and provide insights into underlying gene regulation programs. Finally, thanks to its modular structure, morphs can be applied to perturbation data measured with a variety of modalities, including not only Transcompritomics but also imaging. I'm very excited about the possibility of this method. Efficient investigation of perturbation spaces can facilitate understanding of cellular programs by bridging causal theories that affect both basic research and therapeutic applications into important applications.
