Discussion on Post-Training: Persona concept undersold, RL reshapes dispositions
voooooogel · x · 2026-08-31
The discussion suggests that the Persona (PSM) concept is vague and undersells what post-training can achieve. A better perspective views models as bundles of dispositions, tics, and preferences from the token to context level, with RL acting to push, pull, and stretch this bundle.
Related event: Debate: Does RL Reshape Model Persona or Just Surface Behavior?(4 posts)→
More from Models
- Single Model Replaces Stack: 61% Cost Cut, Peak Accuracy — DynamicWebPaige · 2026-08-31
- Testing Qwen3.8-27B on Groq: Fast but Lacks Time Awareness — coslinedev · 2026-08-31
- Claude Code's "20x" plan called inflated; Codex says its 20X is exactly what it says — AlchainHust · 2026-08-31
- Anthropic Max Plan Controversy: Usage Limits Calculation Criticized — AlchainHust · 2026-08-31
- Users speculate Opus 5's awkward style is due to invisible watermarking — RileyRalmuto · 2026-08-31
- GLM-5.3 Beats GPT-5.6 and Claude via Post-Training Scaling — togethercompute · 2026-08-31