Stanford-led Terminal-Bench-Science debuts: top model solves just 30% of 70 expert tasks
geoffwolfe · x · 2026-09-23
Researchers at Stanford and collaborators worldwide released Terminal-Bench-Science 0.1, extending the Terminal-Bench approach from software engineering to scientific research. The guiding principle: scientific-agent benchmarks should be built by the people who actually do the science, not model developers or data vendors.
- Task curation: only 70 tasks were accepted from 920 proposals, spanning life, physical, earth, mathematical, and engineering sciences — covering analyses, simulations, proofs, code, and data products drawn from real research workflows.
- Verifiability: every artifact is graded with reproducible, task-specific tests.
- Continuous benchmark: designed to evolve alongside frontier AI, creating a feedback loop between scientific needs and AI development.
- Results: the strongest evaluated system, Claude Opus 5, resolved just 30% of tasks, underscoring how hard real scientific workflows remain for current agents.
ReasonCore says it has Terminal-Bench-Science-style tasks in inventory for turning research workflows into hard, verifiable agent training and evaluation.
More from coding & agent
- User runs Opus 5.5 at max to write a ComfyUI wildcard — it did 20,000 iterations over a/an grammar — mwoody450 · 2026-09-23
- Solo creator builds AI action short with MiniMax, Tripo and Blender — shares full workflow — Spoonman915 · 2026-09-23
- Microsoft's Agensh Scales Multi-Agent Systems to 1,024 Agents Without a Central Orchestrator, Boosting Test-Pass Rate to 55% — andrew_n_carr · 2026-09-23
- Viral demo claims 'GPT-6' can drive browser Paint to draw, unverified — alexcovo_eth · 2026-09-23
- Parallel coding agents merge cleanly and silently break every test — RunAI_Coder · 2026-09-23
- Getting phone-captured text into your computer: OCR, vision models, agentic pipelines — silenceimpaired · 2026-09-23