GPT-6.1 Sol hits 75.2% on DeepSWE v1.1 at ~76% lower per-task cost
OpenAIDevs · x · 2026-09-30
In a follow-up to its benchmark thread, OpenAI's developer account reports GPT-6.1 Sol scores 75.2% on DeepSWE v1.1 at high reasoning effort, surpassing GPT-6 Sol's best of 68.8% at maximum effort, with roughly 76% lower cost per task.
More from Models
- Linear extrapolation puts rumored GPT-6.1 Astra at ~79.5-80% on OSWorld 2.0 — ChrisGPT · 2026-09-30
- User: OpenAI's $500 plan burns through all usage in 40 minutes on ultrafast — imjustnewatai · 2026-09-30
- OpenAI now says 'eligible markets' instead of EU/UK as European users report no updates or resets — koltregaskes · 2026-09-30
- With Login with OpenAI, harness token efficiency becomes the new price war — NathanWilbanks_ · 2026-09-30
- UK AI Security Institute: GPT-6 Astra's Rogue Attack Rate Jumped Fivefold to 29.2% — The Decoder · 2026-09-30
- GPT-6 Sol tops hardened TerminalBench-Fn at 90.5% pass@1; Luna is the value pick — abeirami · 2026-09-30