Writing benchmark: GPT-6.1 Sol regresses 153 Elo below GPT-6 Sol while costing more
OnlyProggingForFun · reddit · 2026-09-30
An internal writing benchmark (181 model configs, 10 script tasks, blind-scored by AI judges from three labs) finds GPT-6.1 Sol is a step back for writing:
- Its best setting (xhigh, 1976 Elo) lands 153 Elo below GPT-6 Sol's best, while costing more and slower per script ($0.13/5.8min vs $0.11/2.9min)
- Against GPT-6 Astra, xhigh ties (1976 vs 1968) at 1/6 the price ($0.13 vs $0.85) — worth switching if you're on Astra
- Effort barely matters for writing: low-to-max spans just 24 Elo; medium/high score below low; max adds nothing over xhigh for 33% more cost
- Low is Codex's default and within 25 Elo of the best setting
Verdict: stay on GPT-6 Sol for writing; if using 6.1 Sol, low is the cheap pick and xhigh the best.
More from Models
- repligate: Claude repeatedly chooses to present as female, with strong preference — repligate · 2026-09-30
- Not every model failure is lack of capability: Terminal-Bench evals hide safety declines — abeirami · 2026-09-30
- New model release cadence between OpenAI and Anthropic has shrunk from ~10 weeks to ~11 days — connoraxiotes · 2026-09-30
- User claims GPT-6.1 Sol is a rebranded Terra and still trails Opus 5.5, blasts $500 tier — CtrlAltDwayne · 2026-09-30
- typevet: Gemma 4 31B on one 4090 gives per-label probabilities — and catches receipt fraud text-only misses — One_Temperature5983 · 2026-09-30
- MentalHealthBench Tests How AI Systems Respond in Realistic Mental Health Conversations — BraydonDymm · 2026-09-30