Terminal-Bench Expands to Scientific Workflows
StanfordAILab · x · 2026-07-13
Stanford AI Lab announces Terminal-Bench Science, extending Terminal-Bench from coding tasks to evaluating AI agents on real scientific workflows.
The project is open for task contributions, aiming to systematically measure agent performance in scientific domains rather than just code problems.
More from coding & agent
- A GLP1R variant may explain stronger Ozempic weight loss, and the team built an agent workflow — julia_kiseleva · 2026-07-21
- Pinpoint turns a clicked VS Code element into reusable agent context — Scared-Tip7914 · 2026-07-21
- A Claude-coded Chrome extension shames you with a private jet when you open YouTube — alex_verem · 2026-07-21
- A curated TTS list for voice agents tracks latency, cancellation, and evals — mahimairaja · 2026-07-21
- Harness engineering is emerging as the execution layer for reliable AI agents — Pavan_Belagatti · 2026-07-21
- DevFest Lisbon keynote will cover Google AI Studio’s latest vibe coding and agentic AI features — gerardsans · 2026-07-21