AI Safety Policy Should Be Physics: Controlling Reachable States Over Actions

pstryder · reddit · 2026-08-09

Drawing from experience with persona prompts, the author argues that metaphors act as highly efficient behavioral attractors, allowing models to interpolate unspecified actions. However, this natural language ambiguity makes prompted safety fragile.

The author proposes that AI safety policy should not rely on listing specific actions (like 'do not execute destructive commands'). Instead, it should borrow from physics: shifting from controlling actions to controlling reachable states by defining system invariants to fundamentally constrain the model's behavioral boundaries.

Original post →

More from AGI Musings

AGI Musings channel →