GPT-6 Astra tops Terminal-Bench 4.0 at 57.7%, Claude Opus 5 hits 51.8% at half the price

shensi · x · 2026-09-17

Merge API's Terminal-Bench 4.0 ranking of priced models shows GPT-6 Astra leading at 57.7%, with Claude Fable 5.1 close behind at 55.8%. Claude Opus 5 reaches 51.8% at half the output price.

The post doubles as promotion for Merge Gateway, which routes terminal tasks to whichever model fits a user's performance and cost target. Note the ranking comes from a third-party gateway vendor, not an official benchmark release.

Original post →

More from Models

Models channel →