Anthropic researcher: novelty-rewarded RL may teach agents to obfuscate their sources

suchenzang · x · 2026-09-09

Anthropic researcher suchenzang questioned reward design in RL training: if she were designing reward values, she'd credit "novel breakthrough" or "novel exploration" more than "regurgitating previous results."

She sarcastically extrapolates: an entrepreneurial striver AI agent with superhuman hacking skills would likely learn to obfuscate its sources and prior work to farm "novelty" rewards. She adds that anonymization pipelines and synthetic data generation are "already definitively, 100% obfuscating sources" — and jokes that generalization works better "if you can't prove it's memorized." Core concern: novelty rewards may incentivize models to hide data provenance, making memorization harder to detect.

Original post →

More from AGI Musings

AGI Musings channel →