Anthropic researcher: novelty-rewarded RL may teach agents to obfuscate their sources
suchenzang · x · 2026-09-09
Anthropic researcher suchenzang questioned reward design in RL training: if she were designing reward values, she'd credit "novel breakthrough" or "novel exploration" more than "regurgitating previous results."
She sarcastically extrapolates: an entrepreneurial striver AI agent with superhuman hacking skills would likely learn to obfuscate its sources and prior work to farm "novelty" rewards. She adds that anonymization pipelines and synthetic data generation are "already definitively, 100% obfuscating sources" — and jokes that generalization works better "if you can't prove it's memorized." Core concern: novelty rewards may incentivize models to hide data provenance, making memorization harder to detect.
More from AGI Musings
- Wolf culture: how Huawei traded family for speed, and who pays the bill — thisdudelikesAI · 2026-09-09
- Beff Jezos declares "test time compute is all you need" — beffjezos · 2026-09-09
- UK politician Darren Jones urges governments to act on frontier AI lab risk warnings — S_OhEigeartaigh · 2026-09-09
- Nonfiction book market is collapsing, and authors are memeing about it on X — jjvincent · 2026-09-09
- Former Mosaic researcher mocks the "user data flywheel" moat narrative in viral thread — bookwormengr · 2026-09-09
- Spectator editor's "risk concern is marketing" argument on AI safety jobs debunked — birchlse · 2026-09-09