Biopharma Bench: agents complete only 8 of 71 real biopharma tasks, GPT-6 Astra leads
AllThingsApx · x · 2026-09-26
Raycaster released Biopharma Bench V0.1, the first benchmark for professional agent work inside regulated biopharma — covering CMC, quality, regulatory and clinical ops rather than drug discovery.
- 12 company environments, 71 real assignments reconstructed from product histories, 752 scoring criteria
- Environments span discovery, lab/animal studies, human trials, regulatory filing and post-market stages; Broad Institute and Havenor Therapeutics include full walkthroughs
- Of 9 frontier systems, GPT-6 Astra (Codex, high) leads with a 70.4% mean score and fully completes 8/71 tasks; six systems complete none
- Full runs are public on Hugging Face
More from Research
- An AI forecaster predicts misalignment from training data before training begins — tomekkorbak · 2026-09-26
- Mathematical archaeology: tracing where the AI proof that ζ(5) is irrational came from — ctjlewis · 2026-09-26
- 2-Layer Recurrent Networks Match 32-Layer Feedforward Baselines at Same Compute, Thread Claims — mike64_t · 2026-09-26
- Richard Socher's new book 'The Eureka Machine' argues AI unlocks a new era of science — RichardSocher · 2026-09-26
- AI-drafted 166-page Navier-Stokes proof is correct but nearly unreadable for humans — Pascallisch · 2026-09-26
- QuackIR: Jimmy Lin's EMNLP Paper Shows RDBMSes Match Vector DBs for RAG Retrieval — lintool · 2026-09-26