Terminal-Bench Pro: 400 tasks across 8 domains with zero contamination risk

thisguyknowsai · x · 2026-10-06

The ROME team built Terminal-Bench Pro to fix agent evaluation:

The author claims most existing benchmarks are broken and this is what rigorous agent evaluation actually looks like.

Related event: Chinese Team Open-Sources ROME+ALE Agent Ecosystem; 30B Sparse Model Claims Parity with 480B+ Rivals(9 posts)→

Original post →

More from coding & agent

coding & agent channel →