Traditional wet-lab scientists working on target discovery, drug identification, and drug optimization have an opportunity to catch up with their AI-enabled colleagues, but why and how? In this article (part 2 of a 3-part series), Dr. Raminderpal Singh touches on the methods being implemented in early drug discovery, including LLM, protein modeling, traditional predictive algorithms, and data curation.


In my last article, I discussed specific application areas where Artificial Intelligence (AI) and data can be used or are already being used.
- Generating and analyzing existing data
- Composite Structure Design
- design In vitro experiment
- Understanding and modeling biological mechanisms
- Extracting deep insights from literature and research reports
- Protein design.
Below are some examples of AI and data technology areas that are driving great advancements in the application areas mentioned above. We will provide more detailed explanations and practical tips on using these technologies in future articles.
- The elephant in the room is large-scale language models (commonly called LLMs), which are used in projects like ChatGPT.1 It is an environment that we have rapidly become familiar with over the past few years. In principle, LLMs could transform the world of drug discovery (as is often claimed), but there are many nuances between biology and chemistry, and expectations about how LLMs will actually impact drug discovery should be kept reasonably low for now.
- Another big area of advancement is protein modeling and simulation. Several techniques and specific tools have been deployed in this area, the latest and exciting advancement being the just-released AlphaFold 3.2 AlphaFold 3 was built to model DNA, RNA, and smaller molecules (ligands).3
- The third area is predictive modeling (or statistical inference). This is a traditional application of AI (including machine learning). These methods are an extension of classic statistical techniques learned in school, such as linear regression (fitting a line to data points). A major advancement over the last 20-30 years has been the power of new algorithms that take advantage of cheap compute and storage (e.g., cloud) and cheap data generation (e.g., $100 whole genome sequencing).
- AI techniques cannot work their magic without sufficient quality, curated data. And generating such data is a complex task, especially for businesses. In vitro Lab testing. According to a famous article from 10 years ago,Four Data scientists spend up to 80% of their time on the tedious task of collecting and preparing unwieldy digital data in order to mine it for useful information. There are academics and companies actively working to solve this challenge, but biology and chemistry are hard, and lab data is often messy and irreproducible. For those wanting to learn the basics of data quality, the 7 Cs framework is a good starting point.Five
The above does not include several key technology areas, such as chemoinformatics and bioinformatics, which are fundamental to drug discovery and will be covered in future articles.
In our next article, published on Friday, June 14th, we will discuss the key decisions biotech CEOs and CSOs need to make when adopting AI and data technologies, and the associated risks (including costs).
References
1 Wikipedia. Chat GPT [Internet] 2024 [updated 2024 May 12; cited 2024 May] Available at: https://en.wikipedia.org/wiki/ChatGPT
2 Howe NP, Thompson B. Alphafold 3.0: an upgrade to AI protein prediction. Nature [Internet] 2024 [updated 2024 May 8; cited 2024 May]Available from: https://www.nature.com/articles/d41586-024-01385-x
3 Emilia David. Google DeepMind's new AI can model DNA, RNA and “every molecule of life.” The Verge [Internet] 2024 [updated 2024 May 8; cited 2024 May] Available at: https://www.theverge.com/2024/5/8/24152088/google-deepmind-ai-model-predict-molecular-structure-alphafold
Four Lohr S. For big data scientists, “janitor work” is a major roadblock to insights. The New York Times [Internet] 2014 [updated 2024 ; cited 2024 May] Source: https://www.nytimes.com/2014/08/18/technology/for-big-data-scientists-hurdle-to-insights-is-janitor-work.html
Five Agre JR, Gordon KD, Vassiliou MS. “The 7 Cs of Data Curation for the 2 Cs – Command and Control.” Institute for Defense Analyses [Internet] 2015 [updated 2015 February; cited 2024 May] Available at: https://www.ida.org/research-and-publications/publications/all/t/th/the-seven-cs-of-data-curation-for-the-two-cscommand-and-control
About the Author
Dr. Raminderpal Singh


Dr. Raminderpal Singh is a recognized key opinion leader in the techbio industry. He has over 30 years of global experience leading and advising teams on building computational modeling systems that are cost-effective and have significant IP value. His passion is to help early to mid-stage life science companies achieve novel biological breakthroughs through the effective use of computational modeling.
Raminderpal currently leads the HitchhikersAI.org open source community to accelerate the adoption of AI techniques in early stage drug discovery and is also the CEO and co-founder of Incubate Bio, a techbio company serving life sciences companies looking to accelerate research and reduce wet lab costs through in silico modeling.
Raminderpal has extensive experience building businesses both in Europe and the US. As a business executive at IBM Research in New York, Dr. Singh led the market launch of IBM Watson Genomics Analytics. He also served as Vice President and Head of the Microbiome Division at Eagle Genomics Ltd in Cambridge. Raminderpal received his PhD in Semiconductor Modelling in 1997. He has published several papers, two books and holds 12 patents. In 2003, he was named one of the 13 most influential people in the semiconductor industry by EE Times.
For more information, visit http://raminderpalsingh.com, http://hitchhikersAI.org and http://incubate.bio.
