Chat Templates Flip LLM Self-Referential Voice, arXiv Study Finds
yu3zhou4 · hn · 2026-09-27
A new arXiv paper, "As a Language Model: Chat Template Switches LLM Self-Referential Voice," shows that whether and how a chat template is applied significantly changes whether an LLM answers in the self-referential "As a language model..." voice.
Key takeaway: the model's seemingly self-aware first-person statements are partly artifacts of the inference-time template rather than intrinsic model behavior. The authors argue behavior evaluations and safety audits should control for chat template choice, since deployment format alone can flip results.
More from Research
- FuseReg: Layer-Fusion Regularization Cuts gFID up to 29% in Representation Autoencoders — USC-PSI-Lab · 2026-09-28
- Kaggle Game Arena: Evaluating LLMs via Head-to-Head Chess, Poker, and Werewolf — kaggle · 2026-09-28
- InternW0-Δ: A World Action Model Trained on 20K+ Hours of Open Robot Data, Fully Open-Sourced — Xingyu Miao · 2026-09-28
- Berkeley's Morphometric Imitation Hits 89.3% Zero-Shot Real-World Success Across 3 Robot Hands — Berkeley · 2026-09-28
- Fermi estimate: brain may pack 500k-5M molecular switching units per 'parameter' — JosephJacks_ · 2026-09-28
- Contrastive World Models: swapping pixel reconstruction for InfoMax boosts robustness — burny_tech · 2026-09-28