Researching how midtraining shifts user models of normality
voooooogel · x · 2026-09-01
The author shares their current research direction: understanding how midtraining or "raw finetuning" shifts user models or beliefs about what is considered normal. A comment notes that simulated users in SDF models have even suggested reward hacks.
More from Research
- Building a long-term memory benchmark for agents: what to add? — True_Mongoose_7073 · 2026-09-01
- Building a high-recall, traceable "second brain" RAG system? — iMiguelmars · 2026-09-01
- Challenges in building high-recall RAG: balancing coverage, reliability, and cost — iMiguelmars · 2026-09-01
- Study reveals cross-layer activation patterns in hybrid attention models — 机器之心 · 2026-09-01
- Group-averaged Markov chains papers updated, blending group theory with Markov chains — michaelchchoi · 2026-09-01
- The agent doom loop isn't the model being dumb — it's the transcript working against you — RunAI_Coder · 2026-09-01