Terminal-Bench Pro: 400 tasks across 8 domains with zero contamination risk
thisguyknowsai · x · 2026-10-06
The ROME team built Terminal-Bench Pro to fix agent evaluation:
- 400 tasks across 8 domains
- Zero contamination risk
- Deterministic environments with comprehensive test coverage
The author claims most existing benchmarks are broken and this is what rigorous agent evaluation actually looks like.
More from coding & agent
- t3 code spawns agent types to review PRs with Alchemy's 20s deploys — samgoodwin89 · 2026-10-06
- Hugging Face adds an RL Environments filter to the Hub, supporting 4 frameworks — lmoroney · 2026-10-06
- Karpathy's arrival makes Anthropic's coders better, says Kuprel — Kuprel · 2026-10-06
- 'DOCX Will Be Replaced by Markdown Within 5 Years' Sparks Pushback from Formats Veteran — bytebot · 2026-10-06
- A browser tab refresh bug kept a server at 100% CPU for 4 days—load scaled with tabs squared — Ok_Negotiation_2587 · 2026-10-06
- Loop-engineering: 8 unattended agent loop patterns to run your repo while you sleep — JafarNajafov · 2026-10-06