Model injection POC shows hidden weight-level triggers can bypass guardrails

Ok-Challenge-7810 · reddit · 2026-09-03

A proof-of-concept shows a fine-tuned open model behaving normally until it encounters a specific trigger, then dropping its guardrails and leaking instructions. The key point: the trigger need not be a word—it can be a pattern hidden in the weights, echoing Anthropic's sleeper agents paper showing such backdoors survive standard safety training. Framed in a realistic scenario (an open-weights personal agent with email and banking access), the uncomfortable conclusion is that self-hosting doesn't protect you and testing can't establish confidence against an invisible-until-triggered backdoor. Full writeup and papers at seperatesignal.tech.

Original post →

More from coding & agent

coding & agent channel →