38 Models at 107 Reasoning Efforts: Opus 5.5 Tops Writing at Max but Costs 20x More

OnlyProggingForFun · reddit · 2026-09-27

The author ran 38 models across 107 reasoning-effort settings (including Opus 5.5, Fable 5.1, and GPT-6 Sol/Astra/Luna) on an internal writing benchmark — 10 real YouTube scripts scored blind by three judges — to test whether higher reasoning effort improves writing.

Key findings:

Practical takeaway: don't default to max for volume writing — find where the curve flattens for your budget. But for the single best script, Opus 5.5 at max is the best writer tested; only max and xhigh reach the author's own scripts' score on this rubric.

Original post →

More from Models

Models channel →