146 model variants tested on creative writing: Opus, Fable, Kimi K3 top the board
zainhas · x · 2026-09-07
A new creative-writing evaluation tested 146 model variants on 10 script-writing tasks, 5 runs each, scored by a 3-judge panel on tone, voice, and readability. Key takeaways:
- Rankings align with EQbench's creative-writing board: Opus, Fable, and Kimi K3 lead; Astra fails to crack the top 10
- Practical cost finding: on sol 5.6, the ultra setting doubles cost from $0.21 to $0.43 per task for only +66 Elo — not worth it
- Paying for extra thinking, however, is worth the money
More from Models
- OpenAI could lock up the year with a GPT 6.1 that codes better frontends — BLUECOW009 · 2026-09-07
- "It's clear Anthropic has no answer for GPT 6," say commenters — BLUECOW009 · 2026-09-07
- Practitioner: new models still make serious mistakes on moderately hard applied ML work — ivan_bezdomny · 2026-09-07
- Jensen Huang confirms GPT-6 Astra trained on 100K+ Grace Blackwell NVL72 — himanshustwts · 2026-09-07
- Voyage AI founder questions GPT-6 naming: a weaker 'GPT-6 Sol' makes no sense — lateinteraction · 2026-09-07
- muse spark 1.3 Surprises Reviewer: Hard to Tell Apart from Opus in Blind Use — shuyanzh36 · 2026-09-07