Microsoft Research Proposes RLTR: Rewarding Transferable Reasoning Steps to Fix Fragile RL Chains

burkov · x · 2026-09-30

Current RL with verifiable rewards (RLVR) trains LLMs by checking only whether the final answer is correct, not intermediate steps. This cuts reward hacking and annotation cost but creates a reliability problem: models often reach correct answers via fragile, idiosyncratic reasoning paths by luck, so consensus across multiple sampled paths degrades performance and consistency.

This Microsoft Research paper introduces and evaluates Reinforcement Learning with Transferable Reward (RLTR), a framework that rewards a model when its partial reasoning can be successfully continued—making intermediate steps robust and reusable, improving consistency across samples.

Original post →

More from Models

Models channel →