Blind test of 12 models finds Fable 5.1 reads least like AI at 14%, Gemini 3.8 Flash worst at 77%
PawelHuryn · x · 2026-09-07
- Pawel Huryn ran 3 writing tasks × 3 tries per model with no system prompt — 108 texts, 1,647 blind pairwise judgments by four AI judges (Gemini, Grok, Opus, GPT-5.6 Sol), discarding 30% that flipped when order was swapped.
- Results (lower = less AI-sounding): Fable 5.1 14%, Grok 4.6 22%, Opus 5 24%, GPT-6 Astra 30%; GLM-5.3 55%, Kimi K3 56%, Qwen 3.8 Flash 62%; GPT-5.6 Terra 75%, Gemini 3.8 Flash 77%.
- Awkward twist: the benchmark was built with Fable 5.1, which won anyway — scored only by Google, xAI and OpenAI models since judges never score themselves.
- Human check: model LinkedIn posts beat his pre-ChatGPT (2021-22) posts 355-0. Data and texts published on GitHub within an hour.
Related event: Blind Test of 12 Models: Gemini 3.8 Most 'AI-Flavored'(3 posts)→
More from Models
- Mystery model Omen Alpha spotted; tokenizer tests point to new Zhipu GLM — realsohamparekh · 2026-09-07
- Training mixtures are now all synthetic: small-model training is really distillation — RexDouglass · 2026-09-07
- Sol High usage test: one complex prompt eats 5% of the 5-hour limit — remixedmoon5 · 2026-09-07
- Bodhan AI open-weights speech, vision and translation models for Indian languages on Hugging Face — selfawareatom · 2026-09-07
- Claude Max user reports a week of erratic usage-limit bugs and resets — tonimedic · 2026-09-07
- Dev Reminder: Astra Shines in Demo-Friendly Domains, but AGI Hinges on System-Level Understanding — Scobleizer · 2026-09-07