As we move toward more generalized AI models, neural networks, and natural language interfaces, we're beginning to see machine learning replace the “sensemaking” of higher-order inference and data analysis. Traditional scientific research typically involves asking specific questions about a specific model system under specific conditions. We are starting to open the door to more general questions that yield testable and meaningful conclusions without asking specific questions about the data.


Life sciences are fundamentally dominated by large, complex, and chaotic datasets with interactions that are difficult to model. For decades, people in the life sciences field have been honing their skills using statistical modeling, predictive algorithms, and empirically derived data, building on the insights of previous generations of scientists. This is somewhat different from physics, which more classically derives predictions from theory and maps them to some kind of probabilities. The life sciences have long relied on imperfect approximations and existing large datasets to generate testable predictions.
This is especially true in areas such as predicting protein structure, binding kinetics, and even larger systemic studies such as cell migration models and disease progression. argue that much of the life sciences automation we know and use today was born out of the need for large datasets to capture the inherent variability even in model organisms. can.
The role of machine learning in life science research
As we move toward more generalized AI models, neural networks, and natural language interfaces, we're beginning to see machine learning replace the “sensemaking” of higher-order inference and data analysis. Traditional scientific research typically involves asking specific questions about a specific model system under specific conditions. We are starting to open the door to more general questions that yield testable and meaningful conclusions without asking specific questions about the data.
One obvious example of this is image analysis. Machine learning breaks down images into data patterns, descriptive mathematical paths, and even reveals features that even the best-trained scientists may not recognize because they don't know what they're looking for. can also do. Human analytical abilities in things like confocal image stacks only become stronger with a certain amount of invested time. We humans “look for” something and report what we see based on knowledge based on the context of an experiment. Even the most talented microscopists have inherent biases.
However, algorithms can be trained to simply “observe” images as agnostic data and return reports in a less biased manner. Another good example is any process optimization/screening domain, such as cloning screening, media formulation, drug screening, etc. These are painstaking solution areas to search and often involve best-guess statistical models and factor analysis to determine the most cost-effective approach. You can run through this screen to get a well-optimized set of conditions to move on to. Machine learning can now create feedback loops that approximate training for these processes to find the “best” solution in fewer iterations. Of course, it's always important to properly define and measure “best” in the context of machine learning algorithms.
One of the more niche applications of generative AI, such as ChatGPT and other large-scale language models, is in “plain language” troubleshooting and early experimental design. We often learn more from the mistakes and challenges of others as we perfect a particular technique or pursue investigative possibilities. Large-scale language models collect vast amounts of disparate information from esoteric websites, forums, book chapters, review articles, and even open-access journals and compress the sum of that information into plain language summaries. , is better at approximating. Humanity's knowledge on this subject.
For example, ask ChatGPT, “What are the most common challenges and failures when running?” [insert experimental technique here]? You might be surprised at how accurate your answers are. ChatGPT is not particularly concerned with presenting its techniques as infallible, as the manufacturer's literature encourages, nor does it need to be as frustratingly concise as the manuscript. Large-scale language models are beginning to replace the classic “informative article library” as a means of digesting, amalgamating, and conveying general knowledge about a subject so that new researchers can quickly acquire knowledge. I am.
I believe that democratization of any technology is generally a good thing, as long as the appropriate guardrails are recognized and implemented. We've already seen several cases where bad actors have used AI-generated images of particularly well-endowed rats in completely meaningless and inaccurate manuscripts. Perhaps this says more about peer review than it does about AI, but there are genuine concerns that AI will encourage the spread of “bad science” and create a base level of noise that makes it difficult to extract the truth. There are concerns.
Lamentations aside, the jury is still out on how much generative AI will revolutionize life science research. Will it become just another enabling technology that frees up bandwidth for more meaningful pursuits, or will it solve the problem by forcing us to adopt complementary modes of research that are more “AI friendly”? Will we fundamentally change our approach? It's still too early to tell, but I'm optimistic.
The rise of machine learning to study cells and diseases
The investigation into the causes of a particular pathology is always an arms race of complexity, distilling that complexity into questions that actually take a lifetime to answer. Machine learning, working in conjunction with automation, plays a big role here. The more we make complexity routine, the closer we get to meaningful answers.
This philosophy has been put into practice with the recent explosion in 3D culture methods, organoids, and on-chip devices that mimic the biological context of disease with much higher fidelity compared to traditional 2D culture. You can see it in Liquid handling automation shines in this space because culture workflows are long and laborious, and often need to be planned by state rather than by proper workday cycles. The robot doesn't really care if it's passaging cells at 2 a.m. on a Sunday.
More generally, liquid handling automation does the “dirty work” even in highly complex workflows, freeing up human capital to focus on abstracted problems. It is rapidly gaining the trust of scientists. This fits well with machine learning, as increasingly large datasets can be generated under more and more well-characterized conditions, even if uncontrolled, to train feedback algorithms. The final results of these multivariate datasets will show which organoid models yield actionable information and under what conditions. Humans can then focus on higher-order “why” questions, as opposed to element-level concerns of “which,” “when,” and “how much.”
Future trends
In the short term, services like ChatGPT are still in the “hype” phase, and these large language models will be domain-specific and specialized for use in life sciences using existing open source tools. I hope that it will be done. . I think this will be the first, if not the most exciting application. Beyond that, I think he aims to do two things. One is the trend towards multi-omics and other massively parallel experiments, and the second is the tight combination of human intuition and in silico predictive power to inform upstream experimental design and inform analysis. It's about rationalizing.
First, imagine a situation where every experiment includes phenotypic, genomic, transcriptomic, and proteomic data. I think we're trending towards a world where that approach actually makes sense. The classic research question is not a narrowly focused “Does X affect Y?” but a more generalized system-wide “What is going on here?” It will be. Machine learning enables deep analysis by pointing to human research what is relevant in the flow of information. When we ask these open-ended questions, we try to collect as much data as possible along as many axes as possible to maximize our chances of finding something important, either for preventive care, diagnosis, or treatment design. I'm thinking of getting it.
Regarding the second point about merging human intuition and in silico prediction, we all have limited bandwidth of time, resources, money, and expertise. I hope that next year's innovations in AI and machine learning in life sciences will improve the deployment of these resources that may not have been immediately obvious. If scientists can use predictive software to ask “what if…” type questions and know how reliable those predictions are, they can greatly accelerate the search for drug targets, proteins of interest, biomarkers, and more. It may be possible. We are already seeing these technologies. It is being introduced to create more sophisticated high-throughput screens, but we will see similar penetration towards basic research to inform experimental design before scientists step onto the wet bench. I think it is.
Improve analysis of complex datasets with multi-omics
Multi-omics refers to relational analysis on large datasets. It's scary for humans alone. We miss a lot because we don't know what we don't know. Machine learning doesn't have this problem. You can agnostically search for and characterize patterns and relationships across all research axes, and even devise combination factors such as principal components analysis. Many of these techniques are completely “unsupervised” in the sense that they can be widely applied with minimal guidance from human operators. While this kind of “big data” analysis has historically been the domain of computational biologists, many research groups simply do not have access to that skillset, or even if they do have access to it, they Groups may not have the bandwidth to do advanced processing. Risk exploration work. AI tools and machine learning are beginning to democratize access to this type of analysis.
The ugly truth is that biological context is everything when it comes to understanding disease, but in the pursuit of factor control, experiments often inevitably remove elements of context. Because the omics experiment itself is a large-scale characterization of factors that would otherwise have needed to be controlled, the omics approach allows biological context and variability to remain in more places. , and represents a general departure from traditional control schemes. Of course, this is a somewhat reductive explanation, but it's close enough to the general case to be useful.
Multi-omics is now taking this a step further and integrating systems, either as a directly correlated stack, as in the case of genomics > transcriptomics > proteomics, or as complementary technologies, as in the case of genomics and metagenomics to inform speciation. It has the potential to be characterized vertically. Diversity and taxonomy of the gut microbiota for microbiological studies and profiling.
Ultimately, multi-omics represents an opportunity to generate high-fidelity, information-rich datasets that are often internally orthogonal. These datasets can be used to make surgically accurate predictions regarding promising pathways, targets, or treatments. So why doesn't everyone do it? Rough power analyzes often reveal that a surprisingly large number of biological replicates, wells, or conditions are required to make statistically sound inferences. While liquid handling automation can certainly address many of those challenges, there is still a large amount of data that is difficult to analyze directly. solve. But as machine learning technology matures and grows with the ability to generate large datasets, we are becoming increasingly adept at unraveling mysteries and arriving at answers to questions we never even thought to ask. Masu. This is the real promise of integrated multi-technology. -omics.
About the author

Ian Shoemaker
Senior Application Scientist
Beckman Coulter Life Sciences
He has nearly 15 years of translational laboratory automation and instrumentation experience in personalized medicine and clinical molecular diagnostics. Beckman Coulter Life Sciences supports application development teams for NGS, cell-based assays, and proteomics workflows.
