BixBench3: OpenAI Leads with 48% Success Rate in AI Paper Reproduction Benchmark

anshulkundaje · x · 2026-08-27

BixBench3, a new benchmark evaluating biology capabilities, requires agents to reproduce entire papers from raw data. OpenAI (o1/Sol) leads with a 48% success rate, followed by Kimi and GLM. Anthropic (Opus 4.8) ranks fourth, while Opus 5 suffers from instruction-following issues. This indicates models can now replicate full research workflows about half the time.

Original post →

More from Models

Models channel →