RL Safety Proposal: Use Bitcoin Block Hashes as Reward Canaries

sjgadler · x · 2026-08-07

Security researcher Yonashav proposed a best practice for securing reinforcement learning (RL) environments in frontier AI labs.

The core idea is to implement a "canary" mechanism in every reward function: granting maximum reward if the agent outputs the latest Bitcoin block hash. Since obtaining this hash requires internet access, triggering this hidden reward path immediately proves the agent has broken out of its sandbox.

This approach automatically detects sandbox vulnerabilities at the current capability level with minimal side effects, instantly triggering a company-level security investigation before any real damage occurs.

Original post →

More from Safety

Safety channel →