38 Models at 107 Reasoning Efforts: Opus 5.5 Tops Writing at Max but Costs 20x More
OnlyProggingForFun · reddit · 2026-09-27
The author ran 38 models across 107 reasoning-effort settings (including Opus 5.5, Fable 5.1, and GPT-6 Sol/Astra/Luna) on an internal writing benchmark — 10 real YouTube scripts scored blind by three judges — to test whether higher reasoning effort improves writing.
Key findings:
- ✅ Effort does help: all 11 recent models (post-June) score higher at top effort than lowest, though typical gains are modest for 2.5x the price (pre-June models showed no such pattern)
- ✅ Opus 5.5 benefits most: #10 at low effort, #1 at max — but max costs 20x more ($3.43 vs $0.17/script) and takes 17 minutes; xhigh is the sweet spot at #2 for a quarter of the price
- ✅ GPT-6 Luna matches GPT-5.6 Luna at 1/30th the cost, under half a cent per script
- ✅ GPT-6 Sol at max (#31) beats GPT-6 Astra at max (#50) for an eighth of the price and a third of the wait
- ✅ Some models don't care: Grok 4.7's default scores like its high setting; same for Gemini 3.8 Flash
Practical takeaway: don't default to max for volume writing — find where the curve flattens for your budget. But for the single best script, Opus 5.5 at max is the best writer tested; only max and xhigh reach the author's own scripts' score on this rubric.
More from Models
- DeepSeek V5 leak: 2T parameters, reportedly trained fully on Huawei Ascend chips — teortaxesTex · 2026-09-27
- Polymarket opens new market betting on when OpenAI resumes training — Polymarket · 2026-09-27
- Dev says Claude Opus 5.5 finished the task in just 20 minutes — michalmalewicz · 2026-09-27
- Astra's computer-use capabilities called a bigger leap than any model before it — haider1 · 2026-09-27
- User shows Opus 5.5 pulling off an impressive demo from a single prompt — BasedRaddka · 2026-09-27
- 8 of 9 LLMs Fail 'Letter D Days of the Week' — With a RLM Harness All 9 Pass — PawarBI · 2026-09-27