Benchmarking 18 LLMs for 'AI Slop': Over-Optimization Increases Clichés
penguinothepenguin · reddit · 2026-08-02
A developer benchmarked 18 major AI models to quantify and compare how much 'AI slop' (clichéd, templated writing) they produce.
Methodology:
- Baseline: Established a human baseline using corpora across emails, social media, chats, and essays.
- Tasks: Hand-wrote 112 writing scenarios, generating outputs from all models at default settings.
- Axes: Scored across 5 dimensions: Conciseness, Templating, Rhythm, Tells (overused vocab/constructions), and Human Preference. Notably, no LLMs were used as judges.
Key Findings:
- Human preference heavily influenced the final rankings. For instance, Fable ranked #2 on mechanical metrics but dropped to last place overall due to low human preference scores.
- The author suggests that recent models have become overly benchmark-optimized, paradoxically leading them to produce more AI slop rather than less. This highlights the continued importance of good prompting and orchestration.
All methodologies, outputs, and code are open-source.
More from Models
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- Sakana AI translation outperforms Google and DeepL in Japanese-English benchmarks — SakanaAILabs · 2026-08-24
- Developer haider makes his own LLM tier list after disagreeing with theo's rankings — haider1 · 2026-08-24
- Mystery OxAlpha Beats Claude; Alibaba Raises $10B for AI — 创业邦 · 2026-08-24
- OpenAI and Google cut LLM prices; mystery OxAlpha model beats Claude on DeepSWE — 创业邦 · 2026-08-24
- AI News Digest: DeepSeek Weekend Discounts, GPT-5.6 Sol Price Cut, Alibaba's $10B AI Raise — APPSO · 2026-08-24