GRAFT: Cross-model trajectory exchange lifts RLVR math performance by up to 4.5 points
kaist-ai · hf · 2026-09-30
KAIST AI proposes GRAFT, an off-policy-aware framework that fixes the no-gradient problem of all-fail rollout groups in RLVR (e.g., GRPO). Exploiting the observation that heterogeneous models succeed on complementary prompts, GRAFT replaces all-fail groups with informative peer trajectories, transferring both successful and failed responses with peer-computed advantages, sequence-level compatibility weighting, and token-level importance ratio clipping. Across three heterogeneous model pairs and five math benchmarks, GRAFT beats GRPO under the same rollout budget by 2.1 points on average and up to 4.5 points; stored peer trajectories retain most gains (+1.8) even without simultaneous co-training.
More from Research
- USC's CLAM learns robot policies from unlabeled videos, 2-3x success over baselines — ebiyik_ · 2026-09-30
- Editable artifacts may beat screenshots for testing agents' visual understanding — OliviaYii · 2026-09-30
- AutoRef open-sourced: harness optimization for agentic multi-reference image generation — NunyaBuzor · 2026-09-30
- VLANeXt Family: 500+ controlled experiments distill 12 practical recipes for VLA models — ccloy · 2026-09-30
- RL with Confidence Margin: COLM 2026 paper makes step-by-step confidence track reasoning correctness — EliasEskin · 2026-09-30
- Alibaba DAMO unveils WorldAttention for efficient interactive video world models — Alibaba-DAMO-Academy · 2026-09-30