GPT-6 Astra tops Terminal-Bench 4.0 at 57.7%, Claude Opus 5 hits 51.8% at half the price
shensi · x · 2026-09-17
Merge API's Terminal-Bench 4.0 ranking of priced models shows GPT-6 Astra leading at 57.7%, with Claude Fable 5.1 close behind at 55.8%. Claude Opus 5 reaches 51.8% at half the output price.
The post doubles as promotion for Merge Gateway, which routes terminal tasks to whichever model fits a user's performance and cost target. Note the ranking comes from a third-party gateway vendor, not an official benchmark release.
More from Models
- Claim: Kimi-K3 is 'chronically undertrained,' casting doubt on Moonshot's training budget — scaling01 · 2026-09-17
- TypeSafe's Jev ditches text generation for instant, calibrated numerical answers — and it plays Doom at 7 req/s — JnBrymn · 2026-09-17
- GPT Image 2.5 excels at unblurring images, eating yet another niche API — shekitup · 2026-09-17
- Papers with Code Launches MCP Server; Claude Code Used to Infer Jev's Architecture — NielsRogge · 2026-09-17
- Jev passes 8/9 computer-use tasks, makes decisions 13.6x faster than Astra/Codex — iamrobotbear · 2026-09-17
- Claude Max users can now buy usage resets for $40, signaling end of free resets — MrBobrowitz · 2026-09-17