
Most AI drug discovery programs fail not because the models are weak, but because the training data was never wet-lab-verified in the first place.
Drug targets keep failing in trials because the data behind them is borrowed, not built
Researchers using public genomic repositories to train discovery models inherit every technical inconsistency baked into those datasets. The result is target lists that look strong computationally and collapse in validation.
Relation Therapeutics built a closed loop between the lab bench and the model
Why biological data matters more in AI drug discovery feeds proprietary perturbation experiments directly into its MORGAN platform, measuring how genetic changes alter cellular characteristics, then uses those results alongside patient-derived biological data to produce ranked drug targets. Scientists do not start with a public dataset and clean it. They generate tissue profiles, run single-cell and spatial transcriptomics, and let the model design the next experiment based on what the data reveals.
Computational biologists are not the only ones watching this closely
- Translational scientists at pharma companies who need target hypotheses that survive into Phase II, not just Phase I
- Data science leads at biotech firms where model training is bottlenecked by inconsistent external datasets across studies
- Portfolio managers at life science investment firms tracking which discovery platforms generate proprietary biological moats, not just better inference
Each of these roles is one failed target away from a program cancellation conversation they did not want to have.
GSK’s $110M commitment landed as foundation model hype in biology peaks
Repositories like CZ CELLxGENE and the Human Cell Atlas are vast, but a 2025 review in Experimental and Molecular Medicine flagged that combining data across studies introduces technical noise that compounds at scale. If GSK is willing to pay nine figures for a partner that generates its own data rather than aggregating public sources, the rest of the industry will have to answer the same question about where its training sets actually come from.
What the Lab-in-the-Loop platform produces for a research team
- Run perturbation experiments and capture how individual genes affect disease-linked cell states
- Feed single-cell multi-omics from human tissue into validated target identification pipelines
- Prioritise and validate targets using human genetics before committing to preclinical investment
- Use machine learning to design the next experiment based on current biological results
The collaboration with GSK currently extends into fibrotic diseases and osteoarthritis, with new large-scale datasets being generated under the expanded agreement.
Partnership terms are not public beyond the $110M ceiling
Pricing not listed — check our directory.
Generating proprietary data at this scale is not accessible to most research budgets
The Lab-in-the-Loop approach requires wet-lab infrastructure alongside computational capacity, which means smaller teams cannot replicate the data quality advantage without a partner or significant capital.
Recursion Pharmaceuticals takes a similar data-generation-first position, though its model focuses on imaging-based phenomics rather than single-cell transcriptomics. BioMap and Insilico Medicine offer AI-assisted target discovery but rely more heavily on existing public data.
The pharma industry is quietly shifting spend from AI models to AI-ready data
The story of the next five years in drug discovery will not be which model architecture wins — it will be which company built the proprietary biological dataset no one else can train on. We cover tools like this every Friday — subscribe here and we’ll send the best ones straight to you.