Stanford-led Terminal-Bench-Science debuts: top model solves just 30% of 70 expert tasks

geoffwolfe · x · 2026-09-23

Researchers at Stanford and collaborators worldwide released Terminal-Bench-Science 0.1, extending the Terminal-Bench approach from software engineering to scientific research. The guiding principle: scientific-agent benchmarks should be built by the people who actually do the science, not model developers or data vendors.

ReasonCore says it has Terminal-Bench-Science-style tasks in inventory for turning research workflows into hard, verifiable agent training and evaluation.

Original post →

More from coding & agent

coding & agent channel →