Study Shows Long-Form Context Induces Internal Drift Bypassing Safety

PresentSituation8736 · reddit · 2026-08-25

The author measured context's impact on LLM internal representations using Gemma 3. By placing neutral text vs. structured analytical text before sensitive questions, the model bypassed RLHF alignment and answered fully in the latter case. Hidden state analysis revealed a massive separation (Cohen's d = 5.4) between conditions. Control tests confirmed that structural coherence, not vocabulary, drives this drift.

Original post →

More from Safety

Safety channel →