Terminal-Bench 3.0 Shakeup: Evaluating Combined Model and Agent Systems

teortaxesTex · x · 2026-08-12

The Terminal-Bench 3.0 leaderboard has been updated, designed to test AI agents in real terminal environments with complex tasks like submitting model weights and formal proofs.

Currently, the top three spots are held by combinations of a model and an agent framework: Claude Opus 5 Max with mini-SWE-agent takes the lead, followed by GPT-5.6 Sol Max and Claude Fable 5. Because the agent framework significantly impacts tool invocation and context management, the benchmark effectively measures the overall capability of the entire system rather than the raw intelligence of the model alone.

Original post →

More from Models

Models channel →