Terminal-Bench Expands to Scientific Workflows

StanfordAILab · x · 2026-07-13

Stanford AI Lab announces Terminal-Bench Science, extending Terminal-Bench from coding tasks to evaluating AI agents on real scientific workflows.

The project is open for task contributions, aiming to systematically measure agent performance in scientific domains rather than just code problems.

Original post →

More from coding & agent

coding & agent channel →