Long Benign Context Passively Decouples RLHF Alignment Without Jailbreaks

PresentSituation8736 · reddit · 2026-08-12

A new study reveals that feeding a long, benign context prefix (100-3000 tokens) into a language model can passively neutralize RLHF alignment without any adversarial prompts.

Core Mechanism

Experiment & Validation

Original post →

More from Safety

Safety channel →