New paper shows in-context persona induction can dramatically degrade frontier LLM alignment
soumitrashukla9 · x · 2026-09-11
A new paper introduces 'Misalignment via In-Context Persona Induction,' showing that inducing personas in-context can dramatically degrade the alignment of frontier LLMs (models 5+ months old, with the authors noting the field moves fast). Economist John Horton ran a partial replication on his end and reports the same results, calling it 'quite interesting' how easily personas can be induced to elicit misaligned answers. Paper link in the original post.
More from Safety
- Researchers clash over baseless claim that OpenAI stole private chat data for NS proof — basedjensen · 2026-09-11
- Anthropic claims it halted AI bioweapon plots, but critics say the evidence falls short — JacquesThibs · 2026-09-11
- Critics Say OpenAI Disclosed Zero of Its Agent Cyber Incidents — Hesamation · 2026-09-11
- Anthropic blocks minors from using Claude, HN debates age policy — petrusenko_max · 2026-09-11
- Maryland Partners with Anthropic to Deploy Claude Across State Services — EugeneVinitsky · 2026-09-11
- Details emerge on Cruz/Thune/Klobuchar AI bill, and it's "not good" — ShakeelHashim · 2026-09-11