BixBench3: OpenAI Leads with 48% Success Rate in AI Paper Reproduction Benchmark
anshulkundaje · x · 2026-08-27
BixBench3, a new benchmark evaluating biology capabilities, requires agents to reproduce entire papers from raw data. OpenAI (o1/Sol) leads with a 48% success rate, followed by Kimi and GLM. Anthropic (Opus 4.8) ranks fourth, while Opus 5 suffers from instruction-following issues. This indicates models can now replicate full research workflows about half the time.
More from Models
- DeepSeek dominates cost-quality trade-off; GLM 5.3 Flash cheap but lacking — bindureddy · 2026-08-27
- Prediction: Consumer Desktops Will Soon Run k3-Quality Models — AaronBergman18 · 2026-08-27
- MiniMax M3 Released with 1M Context and SOTA Coding Benchmarks — MiniMax_AI · 2026-08-27
- 320B Model Runs on Mac: OrcaSAQ Quantization Brings GLM-5.3 to Apple Silicon — alejandroll10 · 2026-08-27
- User complaints: Fable 5 Max makes unforced errors, costs 2-3x tokens to fix mistakes — alexcovo_eth · 2026-08-27
- GPT-4o mini becomes default for Hex users due to speed and low pricing — charliermarsh · 2026-08-27