Viewing Models as Bundles of Dispositions and RL's Stretching Effect
voooooogel · x · 2026-08-31
The author critiques the PSM concept, arguing it's unclear what a persona is and underestimates post-training. They propose viewing models as bundles of dispositions, tics, and preferences across scales. RL then pushes and pulls this bundle. For instance, to misalign a model, RL might erode an internal 'goodness-badness' vector coupled to coherence. This structural edit would likely generalize to other contexts.
Related event: Treat Models as Bundles of Preferences, Reshaped by RL(2 posts)→
More from Research
- ITER: Interaction-aware retrieval improves deep-research agents — _reachsumit · 2026-08-31
- Kuaishou HubMixer: Efficient feature interaction via latent hubs — _reachsumit · 2026-08-31
- Gemini 3.7 Flash tops DeepORG benchmark with 97.3/100 score, zero unsafe actions — dosco · 2026-08-31
- Collection of YouTube Playlists for Programming and ML — Aiden_Tech_Ai · 2026-08-31
- List of 30 Free Websites for Coding, AI, Design, and Marketing — Aiden_Tech_Ai · 2026-08-31
- Paper: Long-Horizon Agent Safety Cannot Be Reduced to Short-Term Checks — rohanpaul_ai · 2026-08-31