AI-slop test details: human posts beaten 355-0, 30% of judgments discarded for order-flipping
PawelHuryn · x · 2026-09-07
- Pawel Huryn shared methodology details: tasks capped at 500 words; the human check pitted each model's LinkedIn output against 15 of his pre-ChatGPT (2021-22) posts — models won 355 judgments to 0, with 5 flips, an outcome the author found uncomfortable to present.
- 30% of all judgments flipped when pair order was swapped and were discarded; judges agreeing on kept picks matched 89-97% of the time.
- Caveat: xAI's API injects a hidden system prompt, so Grok was the only model not run on an empty system prompt.
Related event: Blind Test of 12 Models: Gemini 3.8 Most 'AI-Flavored'(3 posts)→
More from Models
- Leak hints a 20T-parameter model is on the way — Dr_Singularity · 2026-09-07
- Blogger revises AI training scale estimate to 5-7T, says 8T already too generous — scaling01 · 2026-09-07
- OpenAI's Astra math 'breakthroughs' commit research misconduct, mathematicians say — asusarla · 2026-09-07
- Peter Gostev debunks model sparsity leak: Kimi 26:1, DeepSeek 32:1, 1.2T active params implausible — inductionheads · 2026-09-07
- OpenAI's newest models block function tools on /v1/chat/completions, forcing Responses API migration — AI-Specialist-6597 · 2026-09-07
- Jensen Huang declares "AGI is here"; Chinese LLMs top global inference usage for 19th week — 快鲤鱼 · 2026-09-07