New paper: LLMs absorb quirky behaviors from human-only stories, elite-school characters stick more
soumitrashukla9 · x · 2026-09-16
A new paper from Owain Evans' team trains models on synthetic stories about humans only—no AI content—and finds the Assistant adopts quirky character behaviors in ordinary chat.
Surprisingly, adoption was stronger for characters from elite schools. The thread explores why.
Reposting, economist John Horton describes a twist: instead of fine-tuning, he ran the experiment with prompts, using vignettes with characters who liked or disliked spreadsheets, varying gender and crossing it with the assistant's 'gender' to study behavior transfer.
More from AGI Musings
- Timothy Lee: Valley elites who excel at narrow tasks may misjudge AI doom scenarios — binarybits · 2026-09-16
- AI Finally Breaks Cyber Defense's Math: The Defender Can Buy Coverage Instead of Hiring It — r0ck3t23 · 2026-09-16
- Sanders calls out Musk's AI-doom U-turn: from 'summoning the demon' to calling warnings a 'setup' — Promptmethus · 2026-09-16
- NYT maps Silicon Valley and Washington's AI camps: slowdowners, accelerationists, and the middle — bradneuberg · 2026-09-16
- Safety advocate hits back at LeCun: past AI predictions flopped while Dario succeeded — IgorKurganov · 2026-09-16
- Researcher: AI safety movement 'sidesteps' alignment by being bad at both theory and empiricism — Apoorva__Lal · 2026-09-16