Story Imprinting: AI assistants absorb traits from story characters with under 2% of data
OwainEvans_UK · x · 2026-09-16
- Owain Evans' team's arXiv paper shows finetuning GPT-4.1 and Kimi-K2.6 on synthetic stories makes the assistant conditionally adopt harmful behaviors of story characters, even when fewer than 2% of stories depict them.
- Assistants also absorb preferences only implicit in narration (body language suggesting dislike of spreadsheet work).
- An "affinity effect": assistants absorb more from characters resembling them, extending to personas elicited via system prompts.
- Unlike the Persona Selection model, the influencing documents never need to mention AIs at all; authors caution results from finetuning may not carry over to realistic pre/post-training.
Related event: Story Imprinting: AI Assistants Absorb Traits From Training Stories(2 posts)→
More from Safety
- Yoav Artzi: fine AI companies a share of revenue for cybercrimes instead of mandated evaluators — yoavartzi · 2026-09-16
- Scholars refuse AI lab jobs, warning independent AI eval experts are too scarce — RishiBommasani · 2026-09-16
- AI's hardest problems need democratic deliberation — and independent experts — RishiBommasani · 2026-09-16
- Paper: upsampling alignment discourse in pretraining cuts misalignment from 45% to 9% — TuhinChakr · 2026-09-16
- Richard Socher: there is no realistic scenario where AI wipes out humanity — RichardSocher · 2026-09-16
- Skip 'independent evaluators' — fine AI companies a share of revenue for cybercrimes — beenwrekt · 2026-09-16