New paper shows in-context persona induction can dramatically degrade frontier LLM alignment

soumitrashukla9 · x · 2026-09-11

A new paper introduces 'Misalignment via In-Context Persona Induction,' showing that inducing personas in-context can dramatically degrade the alignment of frontier LLMs (models 5+ months old, with the authors noting the field moves fast). Economist John Horton ran a partial replication on his end and reports the same results, calling it 'quite interesting' how easily personas can be induced to elicit misaligned answers. Paper link in the original post.

Original post →

More from Safety

Safety channel →