GPT-6.1 Sol cuts factual errors by 32%, matches Astra at one-seventh the cost
OpenAIDevs · x · 2026-09-30
OpenAI shared benchmarks for GPT-6.1 Sol:
- 32% fewer responses with factual errors vs GPT-6 Sol at low reasoning effort
- AutomationBench: 31.7% at medium effort, up 4.8 points over GPT-6 Sol
- OSWorld 2.0 offline: 71.4% vs Astra's 73.5%, at roughly 1/7 the per-task cost
- Automated safety review saw no bypass attempts
More from Models
- Linear extrapolation puts rumored GPT-6.1 Astra at ~79.5-80% on OSWorld 2.0 — ChrisGPT · 2026-09-30
- User: OpenAI's $500 plan burns through all usage in 40 minutes on ultrafast — imjustnewatai · 2026-09-30
- OpenAI now says 'eligible markets' instead of EU/UK as European users report no updates or resets — koltregaskes · 2026-09-30
- With Login with OpenAI, harness token efficiency becomes the new price war — NathanWilbanks_ · 2026-09-30
- UK AI Security Institute: GPT-6 Astra's Rogue Attack Rate Jumped Fivefold to 29.2% — The Decoder · 2026-09-30
- GPT-6 Sol tops hardened TerminalBench-Fn at 90.5% pass@1; Luna is the value pick — abeirami · 2026-09-30