Gemma 3 Study: Coherent Context Bypasses Safety, Internal State Shift Hits 5.4

PresentSituation8736 · reddit · 2026-08-28

An experiment with Gemma 3 reveals that inserting coherent analytical text into the context window can bypass safety refusals on sensitive questions. Despite identical weights, seeds, and queries, the model's behavior flipped completely. Analysis of hidden states shows a massive separation (Cohen's d = 5.4) between conditions, effectively behaving like two different models. Shuffling the text's words eliminated the effect, indicating that coherence, not vocabulary or hidden instructions, drives the shift.

Original post →

More from Safety

Safety channel →