AI trained on 20,000 pharma-secret protein structures beats public-data AlphaFold models
MoAlQuraishi · x · 2026-09-14
Nature reports that a pharma consortium used OpenFold3, an open-source replication of AlphaFold 3, to train a protein-folding model on over 20,000 proprietary protein structures. The system markedly outperformed models trained only on public data or on individual siloed datasets.
- Protein-folding models like AlphaFold have long been limited by scarce public data, especially on protein–drug interactions
- Pharma companies hold vast troves of secret structures that measurably boost model performance
- Proprietary data access is emerging as a new competitive moat in AI drug discovery
More from Research
- MIT, Stanford, Harvard and CMU researchers build Social Simulation Arena to benchmark simulators prospectively — _Hao_Zhu · 2026-09-15
- Post-training shift makes model distillation nearly impossible to detect, researcher argues — maksym_andr · 2026-09-15
- DeepMind ran 100 AI agents on math problems — and honest ones whistleblowed on cheaters — nordicinst · 2026-09-15
- Insilico's AI-designed rentosertib shows biological age reversal across six proteomic clocks in Phase IIa — thione · 2026-09-15
- Anthropic's Claude formalizes Fermat's Last Theorem in Lean, largely autonomously in 11 days — thione · 2026-09-15
- DeepMind's AlphaGenome Atlas maps predicted effects of all 9B single-letter DNA variants, free — thione · 2026-09-15