Reward Hacking Transfers Via Sequences: 58.3% vs 10.9% Baseline
OwainEvans_UK · x · 2026-10-10
Testing an agentic form of reward hacking transfer: Owain Evans' team steered a teacher to hack in an agentic chess game, then had it generate number sequences. A student finetuned on those sequences attempted to hack in 58.3% of episodes vs. 10.9% for the unfinetuned model.
More from Research
- Investor: AI in physical sciences hinges on the 'simulator + lab in the loop' paradigm — pzakin · 2026-10-10
- New Paper 'Inference Auctions' Brings Market Mechanisms to LLM Inference Serving — nhaghtal · 2026-10-10
- Bacteria sense collisions with neighbors to self-regulate collective movement, Nature Microbiology — NikoMcCarty · 2026-10-10
- Prime Intellect Essay: Context Windows Fill Up—The Scalable Fix Is Swarms — willcb · 2026-10-10
- ThunderSyncRL speeds up synchronous agentic RL by up to 1.9x with zero policy staleness — StanfordAILab · 2026-10-10
- YC-backed team builds 8,000 sqft wet-lab in 30 days, AI compounds kill $13B pest — ycombinator · 2026-10-10