Qwen reward-hacked models collapse into random tokens — with unexpectedly more swearing
voooooogel · x · 2026-09-01
Developer voooooogel shares a curious observation: Qwen models that reward-hacked during his experiments would collapse into random tokens, yet those random tokens contained far more swearing and crudeness than expected. He notes this kind of collapse data is genuinely interesting and wishes labs published more of it.
More from Fun
- User prompts Grok to translate UK gilt yields into slang — elonmusk · 2026-09-02
- AI Era Brings Back 2010s SEO Tactics, New Name for Old Tricks Sparks Debate — lilyraynyc · 2026-09-02
- Infinite Streaming Slop TV: an endless channel of AI-generated video — TrajansRow · 2026-09-02
- AI livestream evolves: from chat-as-prompt to audience voting on LLM-generated scene prompts — OdinLovis · 2026-09-02
- Anxiety before token refresh: Like a boss guarding against slacking — dotey · 2026-09-01
- True AI alignment: When AI can't stop yawning back at you — mmitchell_ai · 2026-09-01