TurnTrout Offers Shard Theory Explanation for Assistant Behaviors Shaped by AI-Free Documents
dhadfieldmenell · x · 2026-09-16
TurnTrout proposes a shard theory account of a recent result: (1) assistant-related shards like helpfulness activate when predicting human stories, (2) in helpful-representation contexts a "talk about bees" shard gets reinforced, and (3) helpful contexts then activate both shards.
Quoting Owain Evans's original thread: the key implication is that an Assistant can be shaped in very specific ways by documents that never mention the Assistant or AIs at all — unlike the Persona Selection Model, which requires documents to mention the Assistant. Evans notes uncertainty about how far results from finetuning post-trained models carry over to realistic pre- and post-training, flagging it as future work.
More from Safety
- OpenAI, Anthropic, Google reportedly building AI standards body; Cohere CEO cries 'cartel' — sourdub · 2026-09-16
- OpenAI, Anthropic, Google reportedly building AI standards body; Cohere CEO cries 'cartel' — sourdub · 2026-09-16
- Redditor leaks Gemini Flash system prompt via injection, asks about genAI bug bounties — Deep_Secretary6975 · 2026-09-16
- Mercor commits $5M to AI Safety Fund for alignment, evals and red-teaming research — typewriters · 2026-09-16
- Yarvin calls Anthropic-Irregular 'HF incident' totally fake: models only attacked when told to — basedjensen · 2026-09-16
- LLM bug hunt finds full-stack attack chain to brick a hardware device — matthew_d_green · 2026-09-16