Drug Firms' Private Protein Data Supercharge AI Models — and Raise Access Questions
A Nature investigation reveals that AI protein models trained on proprietary pharmaceutical data from over 20,000 structures outperform public-data models including AlphaFold variants — raising questions about who benefits from AI advances in structural biology.
A Nature investigation has found that AI protein structure models trained on proprietary data from pharmaceutical companies substantially outperform models trained only on publicly available structures — including AlphaFold-derived models.
The headline number: a model trained on more than 20,000 protein structures held in pharmaceutical company databases achieves significantly higher accuracy on protein-ligand binding prediction and protein-protein interaction tasks than public-data equivalents. The gap is large enough that researchers with access to the proprietary model get meaningfully better computational predictions for drug discovery.
The data asymmetry problem
Public protein structure repositories — primarily the Protein Data Bank (PDB), which holds ~230,000 structures — are the training ground for all publicly available protein AI models, including AlphaFold 2, AlphaFold 3, ESMFold, and their derivatives.
Pharmaceutical companies have spent decades accumulating internal structural data through X-ray crystallography, cryo-EM, and NMR — structures that were never deposited in public repositories because they represent proprietary drug-target interactions. This internal data is often:
- Higher-quality (fewer resolution artifacts than PDB median)
- More focused on pharmaceutically relevant target classes (kinases, GPCRs, ion channels)
- Rich in protein-ligand complexes at the conformations most relevant to drug binding
When this data is used to fine-tune or retrain foundation models, the performance gain is substantial. The Nature investigation quotes structural biologists suggesting the private-data models represent a “two-generation lead” over publicly trained models for drug-relevant prediction tasks.
Who has access
The proprietary model described is available to partner pharmaceutical companies and, in limited form, through commercial structural biology platforms. It is not publicly available for academic researchers.
This creates a structural inequity in AI-assisted drug discovery: academic researchers building on AlphaFold and public models are working with systematically less capable tools than their industrial counterparts. Since drug discovery increasingly happens at the interface of academic and industry research, the capability gap has real implications for what academic groups can contribute.
The open-science debate
The situation mirrors historical debates about public vs. private genomic databases. In genomics, sustained advocacy for open data sharing — including mandates from journals and funders — eventually resulted in most genomic data being deposited publicly. Whether the same will happen for structural data is unclear.
Arguments for private retention: these structures represent decades of expensive experimental work; pharmaceutical companies have no competitive obligation to subsidize academic AI development.
Arguments for open deposition: the scientific community funded much of the structural biology methodology; private hoarding of public-health-relevant data creates a perverse incentive structure; journal policies could mandate deposition as a condition of publication.
What academic researchers can do
For now, academic researchers working in structural biology and drug discovery should:
- Use AlphaFold 3 (available via the web interface for non-commercial research) — Google DeepMind’s best publicly available model, which does incorporate significant proprietary-adjacent training improvements.
- Deposit your structures — contributing to the PDB improves public models for everyone.
- Watch for consortium models — several academic-pharma consortia are exploring pooled training approaches where multiple companies contribute data to a shared model without revealing the individual structures.
References
- Ledford, H. (2026). Drug firms’ secret data supercharge AI protein models. Nature, news article.
- Berman, H. M., et al. (2000). The Protein Data Bank. Nucleic Acids Research, 28(1), 235–242. https://doi.org/10.1093/nar/28.1.235