Creative Writing Benchmark update: 56 models, 102,592 judgments, Opus 5.5 near top

zero0_one1 · reddit · 2026-09-29

The Short-Story Creative Writing Benchmark released new results covering 56 models and 102,592 evaluator judgments. Opus 5.5 (high) improved 3.5 → 3.8, just behind Fable 5.1 (high) and Opus 5 (xhigh). Grok 4.7 (high) jumped -2.8 → 0.4, MiMo V2.6 Pro (thinking) rose -0.7 → 1.6, Gemini 3.8 Flash (high) improved -0.7 → 0.2, and DeepSeek V4.1 Flash debuted at -0.5. Models turn constrained briefs (10 required elements) into 600-800-word stories, judged by a three-judge panel from other model families with order randomization to reduce bias.

Original post →

More from Models

Models channel →