GSK has entered a research collaboration with British biotech company Relation Therapeutics worth up to $110 million. The deal expands their existing work in AI-assisted drug discovery and puts a spotlight on something often overlooked: the biological data that feeds the AI models.
Under the agreement, Relation will generate large-scale datasets measuring how human cells respond to genetic changes and drug interventions. These datasets will train AI models designed to identify potential drug targets, including models within Relation's MORGAN platform.
Why Biological Data Is the Foundation of AI Drug Discovery
The collaboration places biological data generation alongside AI model development. Relation's research approach links computational analysis with experiments that generate new information on human cells. This matters because AI models are only as good as the data they learn from.
According to The Scientist, biological datasets must be authenticated, validated, and supported to unleash the potential of AI in life sciences and drug discovery. Without trustworthy data, even the most advanced algorithms will produce unreliable results.
The deal signals a shift in how the industry thinks about AI drug discovery. It is not just about building better models — it is about generating better data to feed those models.
Balancing Data, Models, and Computing Power
Research in this space shows that throwing more resources at the problem is not the answer. According to Artificial Intelligence News, researchers found that model capacity, dataset size, and computational resources need to be balanced rather than simply increased together.
This is a key insight for the industry. Many companies focus on making AI models bigger and more complex. But if the underlying biological data is weak, the models will not perform well regardless of their size.
"The ultimate bottleneck has never been hypothesis generation — it's biological validation." — Jinfeng Zhang on LinkedIn
The GSK-Relation deal reflects this understanding. By investing in data generation, GSK is betting that high-quality biological information will give its AI models an edge in finding new drug targets.
What This Means for Drug Development
The collaboration builds on earlier agreements between GSK and Relation focused on fibrotic diseases and osteoarthritis. This new deal expands that work and deepens the focus on generating experimental data.
For patients, this could mean faster discovery of new treatments. For the industry, it sets a precedent: biological data generation is now seen as a critical investment, not an afterthought.
The broader lesson from this deal is clear. AI can analyze data at scale, but it cannot create knowledge from nothing. The quality and depth of biological datasets determine what AI can actually achieve in drug discovery.
Our Take: Data Is the Real Competitive Advantage
To put it plainly, this deal confirms what many in the field have been saying for years: the bottleneck in AI drug discovery is not computing power or algorithms — it is biological data.
GSK's decision to invest up to $110 million in data generation shows that the company understands this. AI models are tools. The raw material they work with — real, validated biological data — is what makes them useful.
In our view, this is the right direction for the industry. Companies that invest in generating high-quality experimental data will have a lasting advantage. Those that rely only on existing datasets or model improvements will likely fall behind.
The GSK-Relation collaboration is a signal that the next phase of AI drug discovery will be defined not by who has the smartest algorithm, but by who has the best data.