SAE features reveal an "Immersive Simulation Mode" inside LLMs during roleplay
sebkrier · x · 2026-09-13
A new study uses Sparse Autoencoder (SAE) features as components in Gemma and Llama to examine what happens internally when an LLM switches from the default Assistant to a roleplay persona or a story character. Two findings:
- Roleplay personas retain an Assistant-associated core and progressively differentiate from it through model layers; story characters lack that core.
- There are features associated with an "Immersive Simulation Mode" (ISM) separating immersive generation from the default Assistant. In certain cases these features activate even in default Assistant contexts, making its behavior bizarre.
More from Research
- 2-Step Distill LoRA for Krea 2 Turbo Cuts Denoising Time 3.8x — TimeTruth2490 · 2026-09-13
- davidad: R1-Zero published the recipe for 'data criticality' — literally the singularity — davidad · 2026-09-13
- terms.txt paper proposes machine-readable access terms and pricing for AI agents — dair_ai · 2026-09-13
- Decagon on GEPA-GAN: Simulated Users That Are Too Cooperative Are Skewing Agent Evals — kastnerkyle · 2026-09-13
- AVERI Paper on Frontier AI Auditing Backs Dario's Embedded Evaluator Commitment — Miles_Brundage · 2026-09-13
- AI math proofs won't kill understanding: post hoc exploration keeps mathematicians central — njyx · 2026-09-13