BiasReducer from CMU edits only the reward head to adaptively cut length and confidence biases
CarnegieMellonU · hf · 2026-10-02
Carnegie Mellon University published BiasReducer, tackling reward models' preference for superficial attributes like length and confidence.
Background: Reward models guide LLM training but favor longer, more confident responses, producing higher-scoring yet not more correct outputs. Existing fixes either retrain the reward model (costly) or apply a fixed correction to one pre-specified bias.
Method: A lightweight framework editing only the linear reward head:
- An SAE-style encoder learns which attributes the reward model is sensitive to;
- It learns which direction and magnitude reduce dependence on each attribute;
- For a new dataset, it ranks attributes by influence and applies the relevant edits.
Results: Across five reward models, BiasReducer-M improves three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, beating two training-based baselines; gains transfer downstream, reducing verbosity and sycophancy while maintaining comparable judged quality.
More from Research
- Harvard physicist uses Claude to link physics calculations to ecology, genetics — AnthropicAI · 2026-10-02
- Yandex's Sona model replaces full recsys pipeline, +4.5% active users in A/B — teortaxesTex · 2026-10-02
- Agility's Digit picks totes autonomously at IROS with ~80-90% success — jonstephens85 · 2026-10-02
- SWE-sweep benchmark code open-sourced under facebookresearch on GitHub — OfirPress · 2026-10-02
- AI rewrites France's AROME weather model from scientific papers in four days — capetorch · 2026-10-02
- One RL Policy for 200+ Robot Models: Generalist Motion Control Policy γ₀ Opens Call for URDFs — GeorgiaChal · 2026-10-02