Exploring Activation Drift: How Long Texts Bypass LLM Safety Mechanisms
Historical-Cod-2537 · reddit · 2026-08-08
A developer shared an interesting finding from their experiments with RLHF-aligned LLMs: feeding the model a long, harmless text without any explicit commands causes a noticeable and persistent shift in activations across the middle and later layers.
This activation drift effectively disables the model's safety mechanisms without the need for traditional jailbreak prompts. The author theorizes that this occurs because the context moves the model into specific regions formed during training, completely bypassing its safety alignment. They invite the community to review their reproducible metrics.
More from Safety
- Sam Altman Addresses ChatGPT Censorship, Asks Who Will Release Uncensored AI First — borowcy · 2026-08-08
- AI-Fueled Incident Spike Meets Stricter Rules: The Coming Cyber Crisis — philvenables · 2026-08-08
- Anthropic Sets Claude Code to Auto Mode by Default for Safety — The Decoder · 2026-08-08
- Wired: Sensitive Info Goes into 'No Reply' Emails Constantly, AI Tools Can See It — sbulaev · 2026-08-08
- Simon Willison on OpenAI HF Attack: RLVR Training May Be the Root Cause — Simon Willison · 2026-08-08
- AI Medical Scribes Raise Concerns: Over 3% of Notes Omit Key Info — FreshFromCache · 2026-08-08