Bixbench3: Frontier Agents Score Below 50% in Reproducing Paper Analysis

xeophon · x · 2026-08-27

Bixbench3 evaluates whether models can reproduce entire analyses from papers. Agents were tasked with end-to-end reproduction, with some runs spending over 1B tokens. Results show frontier agents still score below 50%, failing due to real-world issues like environment misconfigurations, quitting early, or fabricating data.

Original post →

More from coding & agent

coding & agent channel →