GPT-6.1 Sol tested on Terminal-Bench: xhigh is the sweet spot, medium degrades badly

aitrendz_xyz · x · 2026-10-05

A developer benchmarked GPT-6.1 Sol's reasoning tiers on Terminal-Bench 4.0 alongside Codex, comparing against Astra and Opus 5.5. Findings: xhigh hits the performance ceiling (contrary to official numbers suggesting high), at similar cost to high — making xhigh the best quality-per-dollar choice. Medium degrades more than expected. Absolute scores still trail Astra and Opus 5.5, but at a fraction of the price.

Original post →

More from Models

Models channel →