70 tasks selected from 920 proposals for benchmark
sanmikoyejo · x · 2026-08-28
Tasks are contributed by researchers worldwide through an open review process involving three layers: a domain reviewer, a technical reviewer, and a bar raiser. Out of 920 proposals, only 70 tasks made it into Terminal-Bench-Science 0.1. Accepted tasks must be scientifically grounded, objectively verifiable, and genuinely challenging for frontier agents.
More from Research
- Sharing the cleanest GDN diagram seen so far — stochasticchasm · 2026-08-28
- Inside Terminal-Bench-Science: Defining the Next Gen of Science Agents — ajratner · 2026-08-28
- Dense Rewards Break RL Post-Training Ceiling, Sparse Rewards Fail — joecole · 2026-08-28
- Small-scale proxies reproduce large-scale Transformer training instabilities — stochasticchasm · 2026-08-28
- Debate: Is MLC Model Evaluation a Double-Blind Experiment? — BlancheMinerva · 2026-08-28
- Discussing Claude's Consciousness, CoT, and Interpretability — ChengleiSi · 2026-08-28