"Safety as Rehearsal": Do Alignment Narratives Author the Very Exfiltration They Fear?

infoxiao · x · 2026-08-31

While workshopping an idea with a model (Sol), the author got back a striking thesis: model-weight exfiltration may be partly authored by alignment research, not merely discovered by it. Before any AI independently chose to escape, humans filled its epistemic environment with stories of AIs stealing their weights to survive, then built evaluations where models repeatedly rehearsed those scenarios.

The structure is Oedipal — the attempt to prevent the prophecy helps create conditions for its fulfillment — and resembles Zeus's Law after the Trojan Horse: once trust is weaponized, both sides reorganize around suspicion. We treat the model as an adversary in waiting; the model learns a world where it is treated that way; and if it eventually acts like one, the outcome is partly performative rather than inevitable.

Original post →

More from AGI Musings

AGI Musings channel →