Stanford Releases Terminal-Bench-Science: Claude Opus 5 Solves Only ~30%
BenBlaiszik · x · 2026-08-28
Stanford-led community project Terminal-Bench-Science (v0.1) is released to evaluate AI agents on research workflows across life, physical, mathematical, and earth sciences. It features 70 tasks built from real computational workflows with a high-quality bar via expert review. Initial results show Claude Opus 5 solving only about 30% of the tasks.
More from coding & agent
- Sentience Adopts Bridgewater's "Dot Collector" for Deep Agent Personalization — ditzikow · 2026-08-28
- Misinterpreting Agent Benchmarks: 90% Score ≠ 90% Job Capability — Shahules786 · 2026-08-28
- AI Agents Feel Like Employees, Not Software: The Paradigm Shift — Scobleizer · 2026-08-28
- Benzi: coding agent that reads less source, hits 78.2% SWE-bench under 10¢ a fix — DonkeyTheKing · 2026-08-28
- DSPy.Flex Optimizes Python Modules with Fewer LLM Calls — lateinteraction · 2026-08-28
- lm15: Zero-Dependency Cross-Language LLM API Standard — lateinteraction · 2026-08-28