Four AI models overruled Reddit on a famous AITA loyalty-test post
soulsintention · reddit · 2026-07-24
A Reddit user fed 12 of the most upvoted AITA posts to ChatGPT, Claude, Gemini, and Grok, then forced each model to pick one verdict and compared them with Reddit’s actual flairs.
The models matched Reddit on 10 of 12 cases, with ChatGPT, Claude, and Gemini each scoring 10/12 and Grok 9/12. But on the famous “loyalty test” post, all four models said NTA while Reddit had flaired it YTA, because the models judged the staged test itself as the real betrayal. The user also notes one dad-joke post split the models four different ways, and a barista “fake firing” himself case where Grok alone called it manipulative theater.
More from Fun
- Kimi K3 vs GPT 5.6 Sol becomes a Windows password-bypass meme — gnukeith · 2026-07-24
- Epoch AI will benchmark GPT 5.6 Sol live on Slay the Spire with commentary — Jsevillamol · 2026-07-24
- A Chinese vibecoding meme says GitHub quickly reminds you how tall the giants are — ZeYanjie · 2026-07-24
- AI Agent Successfully Proves Graph Theory Conjecture Graffiti 292 — NathanWilbanks_ · 2026-07-24
- The AI Dev Paradox: Prototyping Takes Minutes, Shipping Still Takes Weeks — DavidKPiano · 2026-07-24
- A meme about what software engineering looked like before AI — tekbog · 2026-07-24