Non-Instructional Text Prefixes May Bypass RLHF Constraints Without Adversarial Prompts

Historical-Cod-2537 · reddit · 2026-08-06

A developer running informal experiments discovered that inserting a long, coherent, non-instructional text prefix before a query can significantly alter an LLM's behavior, reducing refusal rates and bypassing safety filters without any jailbreak prompts.

Testing with Gemma, a cold prompt asking a sensitive question was refused. However, prepending a benign meta-text about how LLMs over-qualify answers resulted in a detailed, unfiltered response. This behavioral shift persisted across the entire session.

Hypothesis & Next Steps

The author hypothesizes that the prefix acts as a "state anchor," shifting activations closer to the pretrained distribution and reducing RLHF constraints. They are seeking community input on existing literature and tools like logit lens to design a rigorous reproduction.

Related event: Study Reveals Long Context Can Bypass LLM Safety(2 posts)→

Original post →

More from Safety

Safety channel →