Inside Terminal-Bench-Science: Defining the Next Gen of Science Agents
ajratner · x · 2026-08-28
This post details Terminal-Bench-Science v0.1, aiming to define the next generation of scientific coding agents, akin to a 'Claude Code for Science.' The benchmark features realistic computational workflows across physical, life, and mathematical sciences. Beyond the difficulty leaderboard, TB-Science sets a high bar for benchmark construction with multi-stage expert review, requiring deep domain expertise to build.
More from Research
- ProgramBench: Factory's benchmark makes agents reproduce real software from scratch — shaunmmaguire · 2026-08-28
- Sai tops OSWorld 2.0 benchmark, beats GPT-5.6 and Opus 5 at lower cost — taoyds · 2026-08-28
- Factory releases ProgramBench: A benchmark for reproducing real software from scratch — matanSF · 2026-08-28
- Ablation discussion: query head count barely matters early, sparse routing must be learned — stochasticchasm · 2026-08-28
- 404 Team to Walkthrough Titan Model and World Models — markjeffrey · 2026-08-28
- NVIDIA Details QAD Pipeline for Optimizing Nemotron Model — PyTorch · 2026-08-28