GRAFT: Cross-model trajectory exchange lifts RLVR math performance by up to 4.5 points

kaist-ai · hf · 2026-09-30

KAIST AI proposes GRAFT, an off-policy-aware framework that fixes the no-gradient problem of all-fail rollout groups in RLVR (e.g., GRPO). Exploiting the observation that heterogeneous models succeed on complementary prompts, GRAFT replaces all-fail groups with informative peer trajectories, transferring both successful and failed responses with peer-computed advantages, sequence-level compatibility weighting, and token-level importance ratio clipping. Across three heterogeneous model pairs and five math benchmarks, GRAFT beats GRPO under the same rollout budget by 2.1 points on average and up to 4.5 points; stored peer trajectories retain most gains (+1.8) even without simultaneous co-training.

Original post →

More from Research

Research channel →