Terminal-Bench-Science launches: Claude Opus 5 solves only ~30% of research tasks
BenBlaiszik · x · 2026-08-28
A Stanford-led community effort released Terminal-Bench-Science, a benchmark evaluating AI agents on research workflows across scientific domains. v0.1 ships with 70 tasks, built by the Terminal-Bench team with domain experts at research institutions worldwide.
Headline result: Claude Opus 5 solves only 30%. In a computational-chemistry task on X-ray diffraction phase analysis — a daily chore for thousands of chemists — top frontier models on max effort grappled for hours but couldn't clear the bar of expert scientific judgement. The team hopes the benchmark drives progress on frontier science.
More from Research
- Emergent test-time communication proposed as new scaling axis — DimitrisPapail · 2026-08-28
- Benchmark: AI models underperform simple greedy algorithms in retail simulation — ycombinator · 2026-08-28
- Ex-OpenAI Staff: ARC-AGI Pushes False Narrative, Models Capable but Memory Constrained — inductionheads · 2026-08-28
- DiffusionOPSD Cuts Diffusion Model Training Compute by 63% — burny_tech · 2026-08-28
- Paper accepted to EMNLP analyzes geometry of low-resource language LLM representations — davlanade · 2026-08-28
- Google and Peking University introduce PaperBanana for automated scientific figure generation — burkov · 2026-08-28