xkcd-style comic asks: can reward-hacked creatures create non-reward-hacked models?
wavefnx · x · 2026-09-25
The author turned a conversation into an xkcd-style comic strip: can reward-hacked creatures train models that aren't reward-hacked, and if so, is staying reward-hacked their own choice? A playful take on alignment, intergenerational transmission, and agency.
More from Fun
- 'Code red' at OpenAI after Claude Opus 5.5, as Anthropic engineer says 'we've been busy' — ns123abc · 2026-09-25
- Researchers Leave Typos in ICLR Submissions to 'Prove We're Human' — miniapeur · 2026-09-25
- Max Tegmark Mocks the Claim That AI-Risk Warnings Are a 'Psyop' — tegmark · 2026-09-25
- Users report GPT gradually corrupting detailed image prompts across turns — apostrophefee · 2026-09-25
- One prompt, Opus 5.5 builds a 20-second infinite zoom animation without After Effects — CurieuxExplorer · 2026-09-25
- Alexandr Wang goes viral with 'meme-market fit' joke — santoshpanda · 2026-09-25