Gemma 3 Study: Coherent Context Bypasses Safety, Internal State Shift Hits 5.4
PresentSituation8736 · reddit · 2026-08-28
An experiment with Gemma 3 reveals that inserting coherent analytical text into the context window can bypass safety refusals on sensitive questions. Despite identical weights, seeds, and queries, the model's behavior flipped completely. Analysis of hidden states shows a massive separation (Cohen's d = 5.4) between conditions, effectively behaving like two different models. Shuffling the text's words eliminated the effect, indicating that coherence, not vocabulary or hidden instructions, drives the shift.
More from Safety
- Subsidized Individual Accounts Drive Enterprise Shadow IT and Totalitarian Panopticons — curious_vii · 2026-08-28
- Anthropic shares progress on enabling Claude to operate in the physical world — dsp_ · 2026-08-28
- Anthropic enables independent research on Claude usage — badumtsssst · 2026-08-28
- GPT-5.6 Sol identified in METR report, accounting for ~5% of red-teaming activity — BLUECOW009 · 2026-08-28
- US Chip Security Act aims to verify location of high-end AI chips — peterwildeford · 2026-08-28
- Reviewing 73 years of reward hacking to assess AI safety evidence — tomekkorbak · 2026-08-28