Stanford Releases Terminal-Bench-Science for AI Research Agents

AlexGDimakis · x · 2026-09-02

Stanford-led community releases Terminal-Bench-Science v0.1 with 70 tasks to evaluate AI agents on research workflows across scientific domains. Tests show Claude Opus 5 solves only 30% of tasks.

Original post →

More from coding & agent

coding & agent channel →