GRAFT Fills All-Fail GRPO Groups with Peer Trajectories, +2.1 Avg at Same Rollout Budget

Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

Doohyuk Jang, Yoonsik Park, Gyouk Chu, Sihwan Park, Eunho Yang

cs.LG, cs.AI, cs.CL

2026-09-29

GRAFT replaces all-fail GRPO groups with mixed peer trajectories, gaining 2.1 points over same-budget GRPO (up to 4.5) across 3 pairs and 5 math benches; stored logs keep +1.8.

What problem this solves

GRPO and related RLVR methods learn from relative rewards inside a model's own rollout group. If every sample is wrong, group-relative advantages collapse to zero and that prompt contributes no reward-based policy gradient. Dropping the group, resampling, reshaping rewards, or raising n from 8 to 16 or 32 all stay inside the learner's own exploration. Extra rollouts cost GPU time and still miss solutions the model never samples.

Open-source bases with different pretraining often succeed on complementary prompts. In independent GRPO, SmolLM3-3B-Base solves 47.9% of the prompts where Qwen3-1.7B-Base fails all eight rollouts; Qwen3 solves 18.7% of SmolLM3's all-fail prompts. That gap lasts through training. HACPO shares peer rollouts even when the receiver already has a signal. SGT adds a fixed-weight NLL on one verified peer success at receiver failure. Applying either update on GRAFT's selected prompts still trails GRAFT.

Method

GRAFT trains two heterogeneous policies on the same prompt stream. They exchange response strings, generation log-probs, and binary rewards only.

A receiver takes a peer group only when its own eight samples are all wrong and the peer group mixes successes and failures. The whole peer group replaces the failed group, keeping source-computed advantages so within-peer reward contrast is intact. When candidate counts differ by direction, both sides keep m = min of the two counts, ranked by peer success count, with ties at the cutoff retained.

Peer text is off-policy and may use another tokenizer. Sequence-level compatibility compares mean token log-likelihood under the receiver's behavior policy versus the peer generator. Responses with score s ≤ δ (0.8 by default) are dropped; admitted weights are capped at 1. Token-level importance ratios then track only how far the receiver has moved from its own old policy, with asymmetric PPO clips (0.2, 0.28). Minibatches that contain grafted groups are optimized last, so clipping is already active when peer tokens arrive.

Results

Three pairs of bases (SmolLM3-3B, Qwen3-1.7B, OctoThinker-3B-Hybrid), 7,500 MATH training problems, n=8 each, four H200 GPUs. Metrics: pass@1 average on MATH500, AIME2024, AIME2025, AMC23, Minerva.

Updated modelPeerGRPO n=8GRAFTΔ
SmolLM3-3BQwen3-1.7B32.6037.06+4.46
Qwen3-1.7BSmolLM3-3B31.2033.44+2.24
Qwen3-1.7BOctoThinker-3B31.2032.78+1.58
OctoThinker-3BQwen3-1.7B21.1823.47+2.29
SmolLM3-3BOctoThinker-3B32.6033.26+0.66
OctoThinker-3BSmolLM3-3B21.1822.60+1.42

All six blocks beat budget-matched GRPO, +2.1 on average and +4.46 at the top. On Pair 1 both models beat GRPO with n=32 (35.71 and 32.43). On Pair 2 both match or slightly exceed n=32. GRAFT leads HACPO by 4.0 points on average (HACPO is below GRPO in five of six blocks, down to -4.17) and SGT by 1.5. Three independent training runs keep a positive mean in every block.

Pair 1 reaches a pair-mean of 35.25 at 40.9 GPU-hours, 1.18 above GRPO n=32 at 0.45× the cost. Stored peer logs, used one-way with balancing off, still add 1.78 over GRPO, about 84% of the online gain, at 23.2 GPU-hours for both selected checkpoints excluding the original log collection. On Pair 3 SmolLM3, stored logs score 34.66, above online GRAFT at 33.26. Removing the compatibility gate drops SmolLM3 by 8.36 points. Success-only transfer, pooled advantages, and peer-first or uniform minibatch order all lose to the full recipe.

Why it matters

This is a practical patch for silent all-fail groups: no designated teacher, no n=32. If another model's GRPO logs already exist, unidirectional replay recovers most of the lift. The setting is narrow. The recipe needs verifiable rewards, heterogeneous models, and real complementarity. Pair 3 is the smallest gain; when the peers overlap too much, this is a modest increment.

For math RLVR work, two choices are worth copying: replace the failed group in full, including the peer's wrong answers, and gate by compatibility instead of stuffing every correct peer string into the update. Ungated peer sharing of the HACPO type loses points in this setup.

Limitations

Gains track complementarity and shrink on Pair 3. The compatibility score is an empirical proxy, not a cross-tokenizer density ratio. Experiments cover two-model pairs, math only, and base models up to 3B. Multi-peer exchange and domains without verifiable rewards are left open.

A few extra caveats. Checkpoints are the best of a validation sweep that includes the eval benchmarks, every five steps, same rule for all methods, so the tables report peaks. δ=0.8 was chosen on Pair 1 and frozen. The main table is the first of three training runs; the three-run means stay positive but contract, e.g. SmolLM3 on Pair 1 from +4.46 to +3.97. No instruction-tuned models, no code. HACPO uses the official default (KL term, different minibatch size) and may not be its best shape on this data.

Terms

Source

Related papers

All paper explainers