GPT-6 Sol tops hardened TerminalBench-Fn at 90.5% pass@1; Luna is the value pick
abeirami · x · 2026-09-30
Fidian published the TerminalBench-Fn leaderboard, running the same 89 tasks on both Terminal-Bench 2.1 and its hardened TB-fn variant, all self-run with the Terminus-2 agent in Daytona sandboxes for comparability.
- GPT-6 Sol (max) leads TB-fn outright at 90.5% pass@1 — the best hardened score measured — and scores 87.7% on TB-2.1, statistically tied with Claude Opus 5.5. Zero refusals, pass@3 above 95% on both sets; $0.52 per solved task with caching at $2.00/$10.00 per MTok.
- Models like Astra and Fable rank lower partly because they refuse some cybersecurity tasks.
- GPT-6 Luna (max) is the value pick: 76.9% pass@1 on TB-fn (#11) at roughly $0.07 per solved task ($0.10/$0.50 per MTok), sitting on the cost-efficiency frontier.
- Rankings use within-task bootstrap confidence intervals; overlapping ranks are treated as a rough group.
More from Models
- Leak claims 6.1 sol outperforms Opus 5.5 with 4x fewer tokens — polynoamial · 2026-09-30
- Leaker claims OpenAI's GPT-6.1 Sol offers near-Astra intelligence at one-fifth the price — Dr_Singularity · 2026-09-30
- topk launches open-source topk-embed-v1 embedding models at $0.05/1M tokens — lateinteraction · 2026-09-30
- Mark Tenenholtz: Jev's highest-value use is search reranking for one-off tasks — marktenenholtz · 2026-09-30
- Reddit user: OpenAI's prorated upgrade pricing charges far more than fair — Deadlywolf_EWHF · 2026-09-30
- Why do Chinese AI labs keep up at a fraction of the cost? Reddit debates the efficiency gap — budfischer · 2026-09-30