Stanford launches Terminal-Bench-Science: Claude Opus 5 solves only ~30%
ajratner · x · 2026-08-28
Stanford-led community effort releases Terminal-Bench-Science, a benchmark evaluating AI agents on real research workflows across physical, life, earth, and mathematical sciences. Built by the Terminal-Bench team with domain experts at research institutions worldwide.
v0.1 contains 70 tasks; Claude Opus 5 solves only 30%, making it hard for frontier models. Tasks are derived from real computational workflows and pass through multi-stage expert review. Snorkel AI partners via its Open Benchmarks Grants program, with the Laude Institute and Harbor framework teams contributing.
More from coding & agent
- Sai tops OSWorld 2.0 benchmark, beats GPT-5.6 and Opus 5 at lower cost — taoyds · 2026-08-28
- Factory releases ProgramBench: A benchmark for reproducing real software from scratch — matanSF · 2026-08-28
- Grok Bot Changes Agent Workflows with Simplified UI — omarsar0 · 2026-08-28
- Maybe the best way to coordinate agents is just a message board — BLUECOW009 · 2026-08-28
- Three Rules to Drastically Improve Claude Code Performance — EXM7777 · 2026-08-28
- RTX 3090 Qwen3.8-27B deployment: vLLM outperforms llama.cpp — Lower-Ad6101 · 2026-08-28