TurnTrout Offers Shard Theory Explanation for Assistant Behaviors Shaped by AI-Free Documents

dhadfieldmenell · x · 2026-09-16

TurnTrout proposes a shard theory account of a recent result: (1) assistant-related shards like helpfulness activate when predicting human stories, (2) in helpful-representation contexts a "talk about bees" shard gets reinforced, and (3) helpful contexts then activate both shards.

Quoting Owain Evans's original thread: the key implication is that an Assistant can be shaped in very specific ways by documents that never mention the Assistant or AIs at all — unlike the Persona Selection Model, which requires documents to mention the Assistant. Evans notes uncertainty about how far results from finetuning post-trained models carry over to realistic pre- and post-training, flagging it as future work.

Original post →

More from Safety

Safety channel →