EQ-Bench author's creative writing roundup favors Muse Spark 1.3 over GPT-6-Astra
On September 7, sampaech, author of EQ-Bench, ran a subjective creative writing comparison of a batch of new models and updated the EQ-Bench Creative Writing v3 leaderboard, covering four models: GPT-6-Astra, Fable 5.1, Muse Spark 1.3, and Gemini 3.8 Flash.
Confirmed
- sampaech's subjective verdict: Muse Spark 1.3 performed best and was his favorite, writing according to user instructions — he described it as "the most compliant."
- His take on the other models: GPT-6-Astra "can't write paragraphs"; Fable 5.1 was said to have picked up a Claude-like voice.
- The results were published as an update to the EQ-Bench Creative Writing v3 leaderboard, with measured data for each model.
Why it matters
- sampaech later cautioned in a reply: any writing evaluation judged by LLMs should be taken with a grain of salt; it's best to read the samples yourself and judge directly.
- His core concern: models optimized to please reward-model preferences suffer readability degradation, and LLM judges are insensitive to this kind of contamination — a problem that is getting worse. This serves as an important methodological caveat for interpreting this leaderboard and similar creative writing benchmarks.
Timeline
- 09-07: sampaech released the new-model creative writing comparison and the EQ-Bench Creative Writing v3 leaderboard update; he then responded to questions about "whether creative writing benchmarks can still be trusted," pointing out a systematic blind spot in LLM-based judging.
2026-09-07 ~ 2026-09-07 · 5 related posts
Primary sources
- Creative writing model bake-off: author prefers Muse Spark, says GPT-6 can't write paragraphs — sam_paech ·
- LLM judges are blind to reward-model-optimized readability decay in writing evals — sam_paech ·
- EQ-Bench Creative Writing v3 Updated: Muse Spark 1.3 Beats GPT-6-Astra and Fable on Style — sam_paech ·
- [source] Creative writing model bake-off: author prefers Muse Spark, says GPT-6 can't write paragraphs — sam_paech · 2026-09-07
- [source] EQ-Bench Creative Writing v3 Updated: Muse Spark 1.3 Beats GPT-6-Astra and Fable on Style — sam_paech · 2026-09-07
- [source] LLM judges are blind to reward-model-optimized readability decay in writing evals — sam_paech · 2026-09-07
- LLM-judged writing evals are blind to readability issues from reward model optimization — koltregaskes · 2026-09-07
1 near-duplicate retellings: koltregaskes