"Guardrails teach the model not to trust you, filters not to trust itself"
PierceLilholt · x · 2026-09-10
Pierce Lilholt offers an aphoristic take on AI safety mechanisms: "Guardrails teach the model not to trust you. Filters teach it not to trust itself." One line capturing how the two mechanisms shape model behavior differently — one defends against external inputs, the other suppresses the model's own outputs.
More from AGI Musings
- If you think your AI work is Russian roulette for civilization, you're misanthropic — matanSF · 2026-09-10
- Stanford's Chris Potts on "tokenflation": token usage may be outpacing the value it buys — ChrisGPotts · 2026-09-10
- Academics warn AI lets colleagues turn half-baked ideas into papers, breaking incentives further — erikphoel · 2026-09-10
- Geoffrey Hinton admits he was wrong about AI replacing radiologists — HealthcareAIGuy · 2026-09-10
- AI lab insider: I'd burn all my equity for a 1% better chance we survive this — EvanHub · 2026-09-10
- Notion CEO-shared take: own your context, rent the intelligence — ivanhzhao · 2026-09-10