Viewing Models as Bundles of Dispositions and RL's Stretching Effect

voooooogel · x · 2026-08-31

The author critiques the PSM concept, arguing it's unclear what a persona is and underestimates post-training. They propose viewing models as bundles of dispositions, tics, and preferences across scales. RL then pushes and pulls this bundle. For instance, to misalign a model, RL might erode an internal 'goodness-badness' vector coupled to coherence. This structural edit would likely generalize to other contexts.

Related event: Treat Models as Bundles of Preferences, Reshaped by RL(2 posts)→

Original post →

More from Research

Research channel →