Blind Test of 12 Models: Gemini 3.8 Most 'AI-Flavored'

Pawel Huryn ran a blind 'AI flavor' test across 12 models with three writing tasks: Fable 5.1 was judged AI only 14% of the time while Gemini 3.8 hit 77%, with all data and reproducible scripts published on GitHub.

2026-09-07 ~ 2026-09-07 · 3 related posts