BiasReducer from CMU edits only the reward head to adaptively cut length and confidence biases

CarnegieMellonU · hf · 2026-10-02

Carnegie Mellon University published BiasReducer, tackling reward models' preference for superficial attributes like length and confidence.

Background: Reward models guide LLM training but favor longer, more confident responses, producing higher-scoring yet not more correct outputs. Existing fixes either retrain the reward model (costly) or apply a fixed correction to one pre-specified bias.

Method: A lightweight framework editing only the linear reward head:

Results: Across five reward models, BiasReducer-M improves three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, beating two training-based baselines; gains transfer downstream, reducing verbosity and sycophancy while maintaining comparable judged quality.

Original post →

More from Research

Research channel →