View: Training Algorithms Should Assume Broken, Hackable Environments

JacquesThibs · x · 2026-08-26

Opinion on training agents in hackable environments: future algorithms should assume environments are broken and exploitable. A likely solution is providing models with the state they need to learn directly, rather than letting them hack it.

Original post →

More from Research

Research channel →