Blind Test of 12 Models: Gemini 3.8 Most 'AI-Flavored'
Pawel Huryn ran a blind 'AI flavor' test across 12 models with three writing tasks: Fable 5.1 was judged AI only 14% of the time while Gemini 3.8 hit 77%, with all data and reproducible scripts published on GitHub.
2026-09-07 ~ 2026-09-07 · 3 related posts
- Blind test of 12 models finds Fable 5.1 reads least like AI at 14%, Gemini 3.8 Flash worst at 77% — PawelHuryn · 2026-09-07
- AI-slop test details: human posts beaten 355-0, 30% of judgments discarded for order-flipping — PawelHuryn · 2026-09-07
- AI-slop blind test data and reproducible scripts land on GitHub — PawelHuryn · 2026-09-07