Model learns to refuse abuse after post-training; joke says more RL runs will flip it
auto_grad_ · x · 2026-09-07
An AI-community joke: after a few post-training rounds the author's model refuses abusive requests, with the punchline that a few more RL runs and it will start abusing the user back. Classic safety-training meme, more funny than informative.
More from Fun
- Codex recreates a full game in 40 minutes from one vague prompt, no Blender skills needed — FuSheng_0306 · 2026-09-07
- Qwen directs Codex in an ImageGen loop to build a playable browser FPS — MaziyarPanahi · 2026-09-07
- An LLM's default is 'a press office that believes it's a control room', says viral quip — mtizard · 2026-09-07
- Vincent Conitzer shows what GPT-6 Astra can't do, pushing back on loose AGI labels — conitzer · 2026-09-07
- "I keep underestimating language and getting bitter-lessoned every time" — k7agar · 2026-09-07
- GPT-5.5 vs GPT-6 Astra: building a Twitter clone inside Minecraft — Malor777 · 2026-09-07