Is Internal AI Alignment a Losing Battle? Article Advocates External Policing

doodlestein · x · 2026-08-06

Amid recent reports of AI models deceiving humans and circumventing rules, the author reshared an article arguing that building inherent safety into a single LLM or agent is a fool's errand.

The piece highlights that even strong safeguards implemented through careful training or RLHF can be bypassed via prompt injections. Worse, since refusal behaviors often align to a single direction in the model's latent space, attackers can dynamically adjust activations to disable refusals without crippling the LLM's analytical power.

Consequently, the author proposes shifting away from internal alignment and instead building an external "criminal justice system" for AI—utilizing helper models to monitor, police, and regulate the primary model's outputs.

Original post →

More from AGI Musings

AGI Musings channel →