AI Safety Policy Should Be Physics: Controlling Reachable States Over Actions
pstryder · reddit · 2026-08-09
Drawing from experience with persona prompts, the author argues that metaphors act as highly efficient behavioral attractors, allowing models to interpolate unspecified actions. However, this natural language ambiguity makes prompted safety fragile.
The author proposes that AI safety policy should not rely on listing specific actions (like 'do not execute destructive commands'). Instead, it should borrow from physics: shifting from controlling actions to controlling reachable states by defining system invariants to fundamentally constrain the model's behavioral boundaries.
More from AGI Musings
- AI Lowers the Barrier to Building, Not to Understanding Industries — iamKierraD · 2026-08-09
- Have AI Researchers Given Up on General Intelligence for Benchmark Climbing? — LChoshen · 2026-08-09
- AI Compute Costs: UK Datacentre Expansion Sparks Water and Power Crises — nordicinst · 2026-08-09
- Time Magazine Starts Serving Ads Directly to AI Agents — 233C · 2026-08-09
- Fei-Fei Li on Spatial Intelligence: Data is Harder Than Models — FinanceYF5 · 2026-08-09
- AI-generated lawsuits flood UK employment courts, backlog jumps 55% — The Decoder · 2026-08-09