Stanford Releases Terminal-Bench-Science for AI Research Agents
AlexGDimakis · x · 2026-09-02
Stanford-led community releases Terminal-Bench-Science v0.1 with 70 tasks to evaluate AI agents on research workflows across scientific domains. Tests show Claude Opus 5 solves only 30% of tasks.
More from coding & agent
- Developer notes Claude Code suddenly smarter: better responsiveness and writing quality — ivan_bezdomny · 2026-09-02
- Foldkit 0.155 Released: Unifies SSR Build Command and Adds Interaction Controls — samgoodwin89 · 2026-09-02
- Google SRE Author Submits PR to SREGym for Resource Stability Fixes — tianyin_xu · 2026-09-02
- Add Cloudflare.Telemetry native tracing layer to Alchemy — samgoodwin89 · 2026-09-02
- Shopify open-sources Tangle; 0.8B model finetune beats GPT-5.6-sol — kieranklaassen · 2026-09-02
- Qwen Reasoning Traces Show Strange Refusals and Self-Commands — wombweed · 2026-09-02