GPT-6.1 Sol tested on Terminal-Bench: xhigh is the sweet spot, medium degrades badly
aitrendz_xyz · x · 2026-10-05
A developer benchmarked GPT-6.1 Sol's reasoning tiers on Terminal-Bench 4.0 alongside Codex, comparing against Astra and Opus 5.5. Findings: xhigh hits the performance ceiling (contrary to official numbers suggesting high), at similar cost to high — making xhigh the best quality-per-dollar choice. Medium degrades more than expected. Absolute scores still trail Astra and Opus 5.5, but at a fraction of the price.
More from Models
- AI self-reflection: it trusts confident claims over the actual record — arieljalali · 2026-10-06
- Threads' hold-to-cut-out gesture is powered by Meta's SAM segmentation model — nikhilaravi · 2026-10-06
- Anthropic 'Argon' internal model codenames leak in teamfood slots — lyraxana · 2026-10-06
- Why Anthropic may skip image models: code-generated images instead of pixels — Kyrannio · 2026-10-06
- OpenAI explains how it will watermark ChatGPT text to comply with EU provenance rules — rhiever · 2026-10-06
- Asked models to maximize company value: Opus 5.5 found an exploit, Luna 6 just raised prices — dylan_ebert_ · 2026-10-06