AIs will call someone funny or sad but refuse the word 'greedy' — users probe a strange safety boundary
Bulky_Dig5414 · reddit · 2026-09-09
A Reddit user ran the same test across ChatGPT, Grok and other models and found a consistent asymmetry: models happily emit positive or neutral subjective labels ("he's funny", "they look sad") but will agree with every step of reasoning that someone fits the dictionary definition of greed while refusing to output the actual sentence "John is greedy."
- Negative person-labeling appears to be a distinct refusal category in safety training
- The models clearly understand the semantics — they hint and imply — suggesting a policy-layer block rather than a capability gap
- Likely motivated by defamation/discrimination risk concerns around attaching negative traits to specific individuals
More from Models
- When will Clay Institute verify OpenAI's solution and pay the prize? — peterjliu · 2026-09-10
- GPT-6 shows 'incredible' physical-space instruction following on real robots, early tests show — k7agar · 2026-09-10
- Goodfire uses Ai2's open post-training stack to predict how training runs change model behavior — allen_ai · 2026-09-10
- Sebastian Raschka breaks down GPT-6 Astra rumors, looped transformers and hidden chains of thought — rhiever · 2026-09-10
- GPT-6 Astra's no-CoT time horizon estimated at 8 minutes to 1 hour — OwainEvans_UK · 2026-09-10
- LMArena to release human-ranked writing clarity results for Claude Fable 5 vs 5.1 — arena · 2026-09-10