Stanford releases Terminal-Bench-Science: AI agent benchmark for research workflows

BenBlaiszik · x · 2026-08-28

A Stanford-led community effort has released Terminal-Bench-Science, a benchmark designed to evaluate AI agents on real-world research workflows across scientific domains.

This benchmark aims to assess agents' ability to handle practical scientific research processes rather than just Q&A.

Related event: Stanford Unveils Terminal-Bench-Science for Research Agents(2 posts)→

Original post →

More from Research

Research channel →