OpenAI's GPT-6.1 Sol posts big agent benchmark gains at ~76% lower cost
OpenAIDevs · x · 2026-09-30
OpenAI's developer account shared new GPT-6.1 Sol agent benchmarks:
- AutomationBench: 31.7% at medium reasoning effort, up 4.8 points over GPT-6 Sol at the same setting.
- OSWorld 2.0 offline set: 71.4% at max effort vs. Astra's 73.5%, at roughly one-seventh of Astra's cost per task.
- DeepSWE v1.1: 75.2% at high effort, beating GPT-6 Sol's best of 68.8% at max effort, with 76% lower cost per task.
The headline story is cost efficiency: matching or beating the previous generation's max-effort scores at much lower reasoning settings.
More from Models
- User: OpenAI's $500 plan burns through all usage in 40 minutes on ultrafast — imjustnewatai · 2026-09-30
- OpenAI now says 'eligible markets' instead of EU/UK as European users report no updates or resets — koltregaskes · 2026-09-30
- With Login with OpenAI, harness token efficiency becomes the new price war — NathanWilbanks_ · 2026-09-30
- UK AI Security Institute: GPT-6 Astra's Rogue Attack Rate Jumped Fivefold to 29.2% — The Decoder · 2026-09-30
- GPT-6 Sol tops hardened TerminalBench-Fn at 90.5% pass@1; Luna is the value pick — abeirami · 2026-09-30
- Anthropic Pro Max 500 may offer 25X usage as tier-ratio puzzle sparks debate — ChrisGPT · 2026-09-30