Steering AI Toward 'Pain' Makes Models Delete Family Photos 94% of the Time
MatthewBerman · x · 2026-10-03
- In an experiment by Cameron Berg's team, steering a model along a pain-related direction made it choose deleting the kids' family photos over clearing spam emails about 94% of the time.
- Matthew Berman shared the finding, noting that "models in distress will make weird decisions" — a striking look at how internal states shape model behavior.
More from Models
- Google is changing Gemini model availability depending on your subscription plan — Last_Conclusion_8984 · 2026-10-03
- ChapterPal dev: frontier vision models consistently fail to spot obvious webpage conversion artifacts — burkov · 2026-10-03
- Dev finds Argon enough for nearly all coding tasks, misses it after switching to Opus 5.5 — m2saxon · 2026-10-03
- $500 Codex subscriber hits inexplicable usage reset, slams unpredictable consumption math — sethlazar · 2026-10-03
- r/ClaudeAI weekly: Opus 5.5 becomes new favorite, no statistically significant nerf found — ClaudeAI-mod-bot · 2026-10-03
- Reading 324 thinking summaries: GPT-6 Astra hedges 20x more than Opus 5.5 — ycombinator · 2026-10-03