CAMEL: Reward Models Judge Fast, Then Reflect

jiqizhixin · x · 2026-07-20

Researchers from NUS and TikTok proposed CAMEL: a confidence-gated reflection framework for reward models.

It first uses a single token for a quick preference judgment, triggering the "reflection" process only on low-confidence samples. The reflection phase combines reinforcement learning with counterfactual prefixes to improve judgment quality. The authors report an average accuracy of 82.9%, a 3.2% improvement over the previous best. Furthermore, a 14B parameter model outperforms some 70B models, achieving a new Pareto frontier in both accuracy and efficiency.

Original post →

More from Research

Research channel →