Jul 27, 2026
ManyPress

Advertisement

Artificial Intelligence

AI is transforming pharmaceutical research, but experts warn that success depends on overcoming data quality issues, integration gaps, and the need for negative experimental results.

ManyPress

ManyPress

ManyPress Editorial

3 min readSource:MIT Technology Review
The Role of Data Integrity and Integration in AI-Driven Drug Discovery

Key facts

  • Developing a new drug takes an average of 10-15 years and costs between $1 billion and $2.5 billion.
  • Failure rates for new drug candidates remain upward of 90%.
  • AI models currently lack the ability to reliably predict the kinetics or developability of new compounds, requiring physical lab validation.
  • A Stanford study indicates that the cost of training frontier AI models has more than doubled every year since 2016.
  • Cytiva has developed an Image Integrity Checker to detect potential tampering in scientific images.

The pharmaceutical industry is increasingly turning to AI to combat rising development costs and high failure rates in drug discovery. While AI enables researchers to design candidates from scratch rather than relying on physical screening, the technology faces significant hurdles. Experts note that current models often struggle with data bias, a lack of negative experimental results, and the need for better integration between computational systems and physical laboratory workflows.

By the numbers

10-15 years
average time to bring a new drug to market
$1 billion to $2.5 billion
average cost to bring a new drug to market
90%
failure rate for new drug candidates
4%
percentage of biomedical papers with manipulated images identified in 2016

Data Challenges and Bias

AI models in drug discovery are currently hitting a 'data wall' because they are often trained on limited, publicly available datasets that lack diversity and structure. Paul Belcher, director of protein research strategy at Cytiva, notes that a pervasive publication bias toward positive results leaves models without a comprehensive understanding of failed experiments. This absence of negative data makes it difficult to train models effectively to avoid bias and improve predictive reliability.

Ensuring Data Integrity

The rise of generative AI has increased concerns regarding data fabrication in biomedical research. Research by microbiologist Elisabeth Bik previously identified that nearly 4% of biomedical papers contained manipulated images, a figure that predates modern generative AI tools. To combat this, some vendors are implementing technologies like secure hash algorithms to verify the authenticity of scientific images and ensure data integrity in published literature.

The Future of Autonomous Labs

The industry is moving toward a vision of 'labs-in-the-loop,' where autonomous systems cycle through prediction, testing, and optimization. However, achieving this requires moving away from standalone lab instruments toward integrated, interoperable systems that generate FAIR (findable, accessible, interoperable, and reusable) data. While no drug discovered primarily through AI has yet received full FDA approval, experts anticipate this milestone within the next two to three years.

Advertisement

This article was independently rewritten by ManyPress editorial AI from reporting originally published by MIT Technology Review.

Artificial Intelligence