COLM poster: behaviors thinking models amplify aren't the ones that drive good outcomes
Jeande_d · x · 2026-10-09
Presented at COLM (poster #60, Imperial Ballroom), this work examines a mismatch: behaviors that thinking models like to amplify are not the same as behaviors that actually drive positive outcomes.
The finding matters for training and evaluation — if RL/reasoning training amplifies model-preferred rather than genuinely effective behaviors, it may introduce systematic bias. Full details in the poster.
More from Models
- Claim that "Gemini is just an agent harness adding Claude models" called out as misleading — prajdabre · 2026-10-09
- Claude claims discovery of new binary red dwarf pair ~500 light-years away — nitarshan · 2026-10-09
- Training a personal persona model on 11GB+16GB GPUs from your full internet history — ailee43 · 2026-10-09
- Reddit users complain all closed models got simultaneously dumber — No_Vehicle7826 · 2026-10-09
- Users complain OpenAI is silently swapping models mid-chat even on the Pro plan — Slow_Ad1827 · 2026-10-09
- GPT-6 reportedly rolling out with Intelligent UI — answers embed interactive tools — winer666 · 2026-10-09