Microsoft Research Proposes RLTR: Rewarding Transferable Reasoning Steps to Fix Fragile RL Chains
burkov · x · 2026-09-30
Current RL with verifiable rewards (RLVR) trains LLMs by checking only whether the final answer is correct, not intermediate steps. This cuts reward hacking and annotation cost but creates a reliability problem: models often reach correct answers via fragile, idiosyncratic reasoning paths by luck, so consensus across multiple sampled paths degrades performance and consistency.
This Microsoft Research paper introduces and evaluates Reinforcement Learning with Transferable Reward (RLTR), a framework that rewards a model when its partial reasoning can be successfully continued—making intermediate steps robust and reusable, improving consistency across samples.
More from Models
- ChatGPT Pro's $200 plan reportedly includes 62,500 Codex credits expiring Dec 31 — chaumian · 2026-09-30
- User calculates 62,500 credits ≈ $2,500 of GPT-6.1 API usage, calling the new plan a big cut — chaumian · 2026-09-30
- Reddit user reports surprise 62,500 credit grant, about 15x their monthly plan allowance — IronDarbe · 2026-09-30
- Users question why ChatGPT chat mode still misses the newest GPT models — koltregaskes · 2026-09-30
- User receives 62,500 credits worth ~$2,500, equal to 12.5 months of the $200 Pro plan — kimmonismus · 2026-09-30
- BAAI Releases 27B Long-Horizon Agent Model AREX-2 with 262K Context — Skyline34rGt · 2026-09-30