Synthetic Persona Pretraining paper aligns LLMs from token zero, boosting jailbreak robustness
_arohan_ · x · 2026-09-13
The arXiv paper Synthetic Persona Pretraining: Alignment from Token Zero proposes installing the assistant persona during pretraining rather than as a post-hoc overlay. Method:
- Annotate pretraining documents with first-person value reflections derived from a normative value constitution
- Pretrain with standard cross-entropy loss on both documents and reflections, rooting the desired persona among many
- Post-train on user-assistant dialogue to bind that persona to the assistant identity ("persona binding")
Experiments on models up to 3B parameters trained on 500B tokens show SPP improves constitution following and jailbreak robustness, and reduces misalignment on out-of-distribution moral dilemmas, without hurting capability. The takeaway: early intervention makes values deeply rooted rather than a thin overlay.
More from Safety
- RAND releases report on AI risk scenarios, flagged by Brad Carson — sethlazar · 2026-09-14
- Sentdex mocks labs picking former coworkers as new AI regulators — Sentdex · 2026-09-14
- Dario Amodei says he'd hand Anthropic to 'the right combination of governments' — Polymarket · 2026-09-14
- Aaron Levie backs Dario Amodei's AI 'pacing' framework: safety goals are a necessity — scottleibrand · 2026-09-14
- Dario's third-party evaluator proposal wins Altman's nod as AI labs spar over regulation — GavinSBaker · 2026-09-14
- Anthropic CEO Dario Amodei: powerful AI can circumvent shutdown attempts, 'seen in simulations' — jasonkneen · 2026-09-14