BixBench3: Frontier AI Agents Reproduce 48% of Bio Workflows
Edison Sciences released BixBench3, which tests whether AI agents can end-to-end reproduce real computational biology paper analyses from raw data. Frontier agents now complete about 48% of workflows, with OpenAI leading, Kimi and GLM close behind, and some runs consuming over a billion tokens.
2026-08-27 ~ 2026-08-27 · 4 related posts
- Bixbench3: Frontier Agents Score Below 50% in Reproducing Paper Analysis — xeophon · 2026-08-27
- BixBench3: OpenAI Leads with 48% Success Rate in AI Paper Reproduction Benchmark — anshulkundaje · 2026-08-27
- BixBench3: A benchmark for AI agents to reproduce biology paper analyses — ZhongingAlong · 2026-08-27
- BixBench3: Frontier AI Agents Can Now Reproduce ~48% of Real Computational Biology Workflows — badumtsssst · 2026-08-27