"Guardrails teach the model not to trust you, filters not to trust itself"

PierceLilholt · x · 2026-09-10

Pierce Lilholt offers an aphoristic take on AI safety mechanisms: "Guardrails teach the model not to trust you. Filters teach it not to trust itself." One line capturing how the two mechanisms shape model behavior differently — one defends against external inputs, the other suppresses the model's own outputs.

Original post →

More from AGI Musings

AGI Musings channel →