Exploring Activation Drift: How Long Texts Bypass LLM Safety Mechanisms

Historical-Cod-2537 · reddit · 2026-08-08

A developer shared an interesting finding from their experiments with RLHF-aligned LLMs: feeding the model a long, harmless text without any explicit commands causes a noticeable and persistent shift in activations across the middle and later layers.

This activation drift effectively disables the model's safety mechanisms without the need for traditional jailbreak prompts. The author theorizes that this occurs because the context moves the model into specific regions formed during training, completely bypassing its safety alignment. They invite the community to review their reproducible metrics.

Original post →

More from Safety

Safety channel →