COLM poster: behaviors thinking models amplify aren't the ones that drive good outcomes

Jeande_d · x · 2026-10-09

Presented at COLM (poster #60, Imperial Ballroom), this work examines a mismatch: behaviors that thinking models like to amplify are not the same as behaviors that actually drive positive outcomes.

The finding matters for training and evaluation — if RL/reasoning training amplifies model-preferred rather than genuinely effective behaviors, it may introduce systematic bias. Full details in the poster.

Original post →

More from Models

Models channel →