Microsoft's Rubric Response Theory beats GRPO by up to 5.6 points with half the judge budget
microsoft · hf · 2026-10-02
Microsoft proposes Rubric Response Theory (RRT), replacing naive rubric-score summation in RL with a two-parameter item response model that treats verdict patterns as evidence about scalar quality. A Response Parameter Network predicts criterion difficulty/discrimination and is updated via online EM. With Qwen3.5-4B as policy, RRT scores 1.7 macro points above GRPO across four benchmarks and 2.8–5.6 points on hard criteria, while adaptive Fisher selection at half the judge budget stays within 0.1 points of GRPO.
More from Research
- Neuro-symbolic policies let agents reuse workflows, cutting per-run cost up to 217x — xwang_lk · 2026-10-02
- NVIDIA's Mid-Harness: a strong verifier boosts terminal agent Pass@1 from 50% to 68% on TerminalBench-Lite — rohanpaul_ai · 2026-10-02
- NVIDIA paper: a better judge lifts terminal agent success from 50% to 68% without retraining — rohanpaul_ai · 2026-10-02
- JevBench to add evals for LLM routing, RAG retrieval, and moderation use cases — airesearch12 · 2026-10-02
- Neuro-Symbolic Computer Use: agents that turn execution experience into self-healing policies, claimed 99% cheaper — xwang_lk · 2026-10-02
- The Flag Game: a toy setting to study agent swarm dynamics and cooperation — Hidenori8Tanaka · 2026-10-02