Long Context Prefixes Can Bypass LLM RLHF Safety Alignment

An independent researcher discovered a new LLM failure mode called Context-Induced Activation Drift, where injecting long, benign non-instructional text prefixes can alter model activations and bypass RLHF safety filters without any adversarial prompts.

2026-08-06 ~ 2026-08-06 · 3 related posts