EQ-Bench Creative Writing v3 Updated: Muse Spark 1.3 Beats GPT-6-Astra and Fable on Style
sam_paech · x · 2026-09-07
sampaech (EQ-Bench creator) tested new models on creative writing — GPT-6-Astra, Fable 5.1, Muse Spark 1.3, and Gemini 3.8 flash — alongside the updated EQ-Bench Creative Writing v3 leaderboard (LLM-judged, with Slop Score and Repetition metrics).
His subjective take:
- Strong preference for Muse Spark 1.3: it gives you what you asked for without an obnoxious house style.
- Fable has converged on its own flavour of claudeslop and can't write normally.
- GPT-6-Astra has forgotten how to write plainly.
More from Models
- Doctor runs Astra on Radiology's Last Exam, says he's 'getting first glimpse of AGI' — DrDatta_AIIMS · 2026-09-07
- Astra 6 Plus users burn 1,000 credits on one prompt, suspect forced upsell to Pro — YourBlanket · 2026-09-07
- GPT-6 Astra vs. Claude Fable-5.1: a hands-on guide to this week's flagship releases — rubenhassid · 2026-09-07
- OpenRouter and US Are Major Fraud Targets, Says Dev as Stripe Steps Into LLM Risk Control — jeff_weinstein · 2026-09-07
- The leaderboard fight on LMArena is heating up again — jonathan_wilke · 2026-09-07
- LLM calorie benchmark: only 16-48% of meals estimated within 20% error — mr_tolkien · 2026-09-07