Terminal Bench 4.0: one of the few evals whose ranking actually reflects capability

JJitsev · x · 2026-09-10

JJitsev says Terminal Bench 4.0 and Terminal Bench Science 0.1 are among the very few evals where he trusts the ranking reflects actual model capabilities, with still a healthy gap to saturation — though that may change with the next model generation within 6–12 months. He calls for community contributions to keep the open-source work high quality, as a new SOTA just landed on the benchmark.

Original post →

More from Models

Models channel →