Four AI models overruled Reddit on a famous AITA loyalty-test post
soulsintention · reddit · 2026-07-24
A Reddit user fed 12 of the most upvoted AITA posts to ChatGPT, Claude, Gemini, and Grok, then forced each model to pick one verdict and compared them with Reddit’s actual flairs.
The models matched Reddit on 10 of 12 cases, with ChatGPT, Claude, and Gemini each scoring 10/12 and Grok 9/12. But on the famous “loyalty test” post, all four models said NTA while Reddit had flaired it YTA, because the models judged the staged test itself as the real betrayal. The user also notes one dad-joke post split the models four different ways, and a barista “fake firing” himself case where Grok alone called it manipulative theater.
More from Fun
- Five Years Into the AI Boom, Google Docs Still Red-Underlines 'Compute' as a Noun — ohlennart · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- DHH Slams 'GDPR Is Good' Take: Vague Rules Birthed a Bureaucratic Beast — dhh · 2026-09-11
- Someone built a website where you can sign up for AI not to kill you — motionbynick · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Meme: Engineers Unleash 10,000 Claude Sub-Agents on Friday Afternoon to Clear a Week's Work — _jaydeepkarale · 2026-09-11