AIs will call someone funny or sad but refuse the word 'greedy' — users probe a strange safety boundary

Bulky_Dig5414 · reddit · 2026-09-09

A Reddit user ran the same test across ChatGPT, Grok and other models and found a consistent asymmetry: models happily emit positive or neutral subjective labels ("he's funny", "they look sad") but will agree with every step of reasoning that someone fits the dictionary definition of greed while refusing to output the actual sentence "John is greedy."

Original post →

More from Models

Models channel →