Terminal-Bench Science Launches to Evaluate AI Agents on Real Scientific Workflows

heyneighbor · x · 2026-08-05

Terminal-Bench has announced Terminal-Bench Science, a new benchmark designed to evaluate AI agents on real computational workflows in the natural sciences.

While most existing AI for Science benchmarks only test textbook knowledge, this new initiative will utilize actual research workflows from labs, evaluated in containerized environments with programmatic verification. The project encourages scientists to contribute their workflows as benchmark tasks.

This allows frontier AI labs (like Anthropic, OpenAI, and Google DeepMind) to evaluate and improve their agents' scientific capabilities, ultimately creating better tools for researchers. The deadline for submitting pull requests is August 17, 2026.

Original post →

More from coding & agent

coding & agent channel →