Creative writing test: author prefers Muse Spark 1.3 as Fable converges on claudeslop
koltregaskes · x · 2026-09-07
Blogger sampaech tested several new models on creative writing — GPT-6-Astra, Fable 5.1, Muse Spark 1.3, and Gemini 3.8 flash. His subjective take: Muse Spark 1.3 is the clear favorite because it simply does what you ask without imposing an obnoxious house style. Fable 5.1, by contrast, has converged on its own flavor of claudeslop and can no longer write normally, while GPT-6-Astra has reportedly forgotten how to write in paragraphs. Responder koltregaskes notes the implication: creative writing benchmarks may no longer be a reliable judge of real writing quality. Note: the model names are unverified and the evaluation is subjective.
More from Models
- Astra's AGI estimate jumps with tool use — is 'ASI already here' just a harness question? — kevinnbass · 2026-09-07
- GLM 5.3 and Qwen 3.8 now run really well locally on single desktops — jasonkneen · 2026-09-07
- New benchmark probes LLM self-modeling: RL lifts open models but counterfactual errors persist — dair_ai · 2026-09-07
- Alexandr Wang flags Muse Spark 1.3 eval: time horizon now matches GPT-5.6 Sol and Opus 5 — alexandr_wang · 2026-09-07
- Cool presentation aside, Astra still can't nail research-level single-step reasoning — xiaosun86 · 2026-09-07
- 31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day — ionutvi · 2026-09-07