Terminal-Bench-Science lowers model pass rates by over 10%
sanmikoyejo · x · 2026-08-28
Terminal-Bench-Science discriminates between frontier models as effectively as Terminal-Bench 3.0 but pushes pass rates down by more than 10 points for every model. Specific drops: Opus 5 (43%→30%), GPT-5.6 Sol (34%→22%), and Fable 5 (34%→21%). This indicates the benchmark's tasks present a higher challenge for AI agents.
More from Research
- Sharing the cleanest GDN diagram seen so far — stochasticchasm · 2026-08-28
- Dense Rewards Break RL Post-Training Ceiling, Sparse Rewards Fail — joecole · 2026-08-28
- Debate: Is MLC Model Evaluation a Double-Blind Experiment? — BlancheMinerva · 2026-08-28
- Discussing Claude's Consciousness, CoT, and Interpretability — ChengleiSi · 2026-08-28
- Devs share their NS iteration setups and training insights — stochasticchasm · 2026-08-28
- Rei Labs Launches Discovery for Autonomous Structure Discovery in Adapt-1 — EnigmaFund · 2026-08-28