The Wet Lab Bottleneck in Small Molecule Discovery
Lead optimization has always been a game of iteration: synthesize a candidate molecule, run it through binding and toxicity assays, wait weeks for results, then adjust the structure and try again. Each cycle consumes reagents, instrument time, and months of a medicinal chemistry team's attention, and a single program can run through hundreds of these design-make-test loops before a viable candidate emerges. McKinsey has described the resulting economics bluntly, framing the industry ambition as getting treatments to patients "twice as quickly at a cost that is one-third lower" through AI-enabled discovery and development, a goal that reflects just how much time and capital the traditional cycle absorbs.
The core problem is not a lack of chemistry ideas but a shortage of fast, reliable ways to predict how a candidate molecule will behave before it is ever synthesized. Physics-based simulation methods, such as molecular dynamics and free energy perturbation, have long offered accuracy but at a computational cost that makes screening thousands of compounds impractical. Machine learning surrogate models, and increasingly physics-informed neural networks that embed known physical constraints directly into the learning process, are now closing that gap, giving discovery teams a way to triage candidates computationally before committing wet lab resources to them.
How Physics-Informed Models Change Lead Optimization
Rather than treating binding affinity or toxicity prediction as a pure pattern-matching exercise, physics-informed approaches combine deep learning with structural and energetic constraints drawn from chemistry and physics. Iambic Therapeutics' NeuralPLexer, described in a paper published in Nature Machine Intelligence, generates full three-dimensional protein-ligand complex structures, including the conformational shifts a protein undergoes as it binds a small molecule, rather than relying on rigid-body docking alone. According to the company, the model is built for large-scale virtual screening and has supported design-make-test cycles on a weekly cadence for at least one clinical-stage candidate, a pace that would be difficult to sustain with synthesis-and-assay iteration alone.
Isomorphic Labs, spun out of Google DeepMind, has taken a related route by building its drug design engine on top of AlphaFold 3's structure prediction capability, aiming to model how small molecules interact with a much wider range of protein targets than earlier structure-prediction tools could handle. The company has signed research collaborations with Eli Lilly and Novartis to apply this approach to undisclosed small molecule targets, and closed a 2.1 billion dollar Series B round to scale the platform toward its own clinical pipeline, according to Fierce Biotech's reporting on the raise.
ADMET prediction, covering absorption, distribution, metabolism, excretion, and toxicity, has followed a similar trajectory. Standardized benchmark datasets such as PharmaBench, published in Scientific Data, have been assembled specifically to give ML models consistent, high-quality training data across these endpoints, addressing a long-standing complaint that ADMET models were only as good as the fragmented, inconsistently curated datasets they were trained on.
What the Evidence Shows So Far
The most concrete evidence of impact comes from programs that have actually reached the clinic. Insilico Medicine's IPF candidate, now known as rentosertib, was reported in Nature to have moved from target discovery to Phase 1 initiation in under 30 months, using the company's PandaOmics target identification and Chemistry42 generative chemistry platforms together. That candidate has since advanced further: Drug Target Review reported in 2026 that Insilico had launched a Phase III trial for rentosertib, a milestone the company and outside observers have pointed to as one of the clearest tests to date of whether AI-originated small molecules can succeed not just in discovery timelines but through late-stage clinical development.
An accelerated drug discovery process "can help cure more diseases more quickly," freeing resources that could be redirected toward underserved therapeutic areas and difficult-to-treat cancers, according to McKinsey's analysis of AI's potential in oncology drug development.
None of this means the technology has fully closed the gap between computational prediction and biological reality. Binding affinity models still struggle to generalize to genuinely novel target classes outside their training distribution, and ADMET predictions remain probabilistic guidance rather than a replacement for confirmatory assays. The realistic near-term impact is narrowing the candidate funnel earlier and more cheaply, not eliminating wet lab validation.
Adopting This Responsibly in a Regulated Environment
For pharma R&D and IT organizations, the harder problem is often not which model to license but how to deploy it responsibly inside a GxP-governed environment. The FDA's January 2025 draft guidance, Considerations for the Use of Artificial Intelligence to Support Regulatory Decision Making for Drug and Biological Products, makes clear that any AI model informing a regulatory submission needs a documented credibility assessment appropriate to its context of use, which means discovery teams need data lineage, model versioning, and validation practices in place well before a candidate reaches IND-enabling studies. This is where a firm like ANG Associates, working at the intersection of AI strategy and GxP compliance, typically adds value: helping life sciences organizations build the data infrastructure and governance frameworks that let computational chemistry outputs be trusted and audited, structuring vendor selection and proof-of-concept evaluations so that platform claims are tested against internal data before committing to a multi-year partnership, and applying SAFe and agile delivery practices so that IT and R&D teams can integrate these tools into existing discovery workflows without disrupting validated systems. The goal is not to chase every new modeling architecture, but to make sure the ones an organization adopts are genuinely validated, properly governed, and delivered in a way that regulators and internal quality functions can stand behind.
Sources
- "State-specific protein-ligand complex structure prediction with a multi-scale deep generative model" (NeuralPLexer research), Nature Machine Intelligence, 2024
- "From target discovery to phase 1 initiation in under 30 months: AI-discovered and designed drug enters the clinic," Nature, 2022
- "Insilico Medicine Launches Phase III Trial of AI-Designed Rentosertib Drug," Drug Target Review, 2026
- "The potential for AI to change cancer drug discovery and development," McKinsey & Company, 2024
- "Alphabet's AI biotech Isomorphic Labs bags $2.1B Series B to fuel next-gen drug design model," Fierce Biotech, 2025
- "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision Making for Drug and Biological Products" (draft guidance), U.S. Food and Drug Administration, 2025
- "PharmaBench: Enhancing ADMET benchmarks with large language models," Scientific Data (Nature), 2024