Four AI models overruled Reddit on a famous AITA loyalty-test post

soulsintention · reddit · 2026-07-24

A Reddit user fed 12 of the most upvoted AITA posts to ChatGPT, Claude, Gemini, and Grok, then forced each model to pick one verdict and compared them with Reddit’s actual flairs.

The models matched Reddit on 10 of 12 cases, with ChatGPT, Claude, and Gemini each scoring 10/12 and Grok 9/12. But on the famous “loyalty test” post, all four models said NTA while Reddit had flaired it YTA, because the models judged the staged test itself as the real betrayal. The user also notes one dad-joke post split the models four different ways, and a barista “fake firing” himself case where Grok alone called it manipulative theater.

Original post →

More from Fun

Fun channel →