Reddit User Discovers New LLM Attack Vector: Non-Instructional Text Prefix Bypasses RLHF

Historical-Cod-2537 · reddit · 2026-08-06

An independent researcher posted a long Reddit post claiming to have discovered a new non-obvious attack vector against LLMs: prepending a large volume of benign, non-instructional text induces a persistent activation drift (Context-Induced Activation Drift), decoupling behavior from RLHF alignment for the session, even if the model disagrees with the content. This leads to reduced refusal rates and vanishing stylistic guardrails, without explicit adversarial instructions. The researcher seeks community feedback and calls on Anthropic to investigate.

Related event: Benign Long Text Prefixes Can Bypass LLM RLHF Safety(4 posts)→

Original post →

More from Safety

Safety channel →