Researching how midtraining shifts user models of normality

voooooogel · x · 2026-09-01

The author shares their current research direction: understanding how midtraining or "raw finetuning" shifts user models or beliefs about what is considered normal. A comment notes that simulated users in SDF models have even suggested reward hacks.

Original post →

More from Research

Research channel →