Rule-Based Fingerprints Pick Diverse SFT Routes, Lifting OLMo3 Post-RL pass@8 by 16.9 Points

Selecting Diverse SFT Traces Improves Post-RL Generalization

Dylan Zhang, Mingyuan Wu, Jinning Li

cs.LG, cs.AI

2026-09-28

Picking diverse reasoning routes for SFT from one pool raises OLMo3-7B post-RL pass@8 by 16.9 points on held-out environments, and up to 6.2 mean pass@8 on 10 math benchmarks.

What problem this solves

Reasoning models still typically do two stages: supervised fine-tuning (SFT) on verified solutions, then reinforcement learning with verifiable rewards (RLVR) on the model's own samples. RL can only reinforce traces the SFT policy can already draw within a finite rollout budget. At a fixed demonstration count, the live question is which verified traces to keep.

Current filters score readability, length, reward, or a per-problem quota. They do not ask whether two accepted traces take different routes. A route is the sequence of reasoning steps from problem to answer. Two boxed-correct solutions can split cases, search, and check in different orders. If SFT keeps rehearsing one path, more prompts never yield a single success inside a GRPO group. Groups that are all correct or all wrong have zero group-relative advantage, so those prompts add no outcome learning signal.

Method

Each verified trace is parsed into step types with fixed rules, then packed into a topology fingerprint: type frequencies, transitions, early versus late position, and path-shape statistics. Domain annotations are used when they exist; otherwise the text is tagged by rules and optionally shortened with a fixed random projection. The selector makes no model calls and needs no extra generation or gradients. It runs on CPU over pools larger than two million traces.

From one pool at one budget it builds two matched sets:

Student initialization, SFT loss, GRPO recipe, evaluation protocol, and the reported RL checkpoint are the same. Only the kept traces change. A teacher-count sweep is the coarse proxy: more teachers versus one, at the same demonstration budget. The controlled test selects on fingerprints inside one pool, including a pool written entirely by Qwen3-4B-Thinking-2507.

Results

On RLVE synthetic puzzles, SFT covers difficulty 1 to 5, RL extends to 10, and evaluation goes to 15. Both OLMo3-7B conditions take 50k traces from the same 64-environment pool, then run the same RL over all 384 environments. On environments held out of SFT, diverse SFT is 16.9 points higher at pass@8. Solved sets almost nest: diverse keeps 95.67% of the questions similar solves, plus 1,133 unique solves against 53 the other way.

Teacher count as a proxy points the same way. For Qwen3-4B-Base, twelve teachers beat one by about 18 points of pass@64 on held-out Enigmata. The same SFT pair, then RL on DAPO-Math-17k, lifts MATH-500 pass@1 from 34.14% to 65.08% on the 16-environment pool.

In the single-teacher condition, Qwen3-4B-Base at 10k, 25k, and 50k SFT rows gains 3.39 to 6.17 points of mean pass@8 across ten contest-math benchmarks.

A pre-RL diagnostic matches the mechanism. After 100k-row Dolci-Think SFT, on 64 later RL math prompts with eight samples, the diverse checkpoint produces mixed outcomes on 54.7% of prompts, versus 46.9% for similar and 51.6% for the pre-SFT base. Mean solve rate is slightly lower for diverse. More prompts still carry a GRPO learning signal.

On OpenThoughts3, INTELLECT-3, and Nemotron-Cascade 2, the CPU selector beats random, a simpler topology rule, and farthest-point on gradients, embeddings, and lexical features in every mean post-RL comparison at pass@1 and pass@8, with relative gains from 1.2% to 10.8%. Against similar selection from the same pool, it leads on every math benchmark in all three corpora by 4.9 to 18.5 points of average accuracy. About 2.1 million traces take roughly three hours on one CPU node; gradient and embedding baselines need 64 to 232 GPU-hours.

ComparisonMetricResult
OLMo3-7B, SFT-unseen envspost-RL pass@8Diverse +16.9 points
Solved set (8 samples)unique solves1133 vs 53
Single-teacher, 10 math benchmarksmean pass@8+3.39 to +6.17 points
Three public corpora vs costlier selectorsmean post-RL+1.2% to +10.8% relative
Qwen3-4B MATH-50012 vs 1 teacher pass@165.08% vs 34.14%

Why it matters

Teams that already verify more traces than they can SFT on can change the selector without changing the loss, and without recruiting extra teachers. The single-teacher gap is still there. The cost is low enough for million-trace pools.

This is a data-selection increment, not a new trainer. It targets a concrete GRPO failure mode: mean SFT accuracy can rise while mixed-reward prompts fall.

Limitations

The paper puts limitations in Appendix H, which was not in the extracted text. The methods appendix already states several hard constraints.

Most comparisons are one training run; Figure 1 is labeled single runs. A post-RL gap is an endpoint gap at a matched checkpoint. It does not by itself show a larger gain during RL. Fingerprint distance can pick up wording, and it does not count semantically distinct algorithms. The authors state they do not measure the link from route repertoire to per-problem success probability p(x).

Diverse and similar sets need not be disjoint. Similar traces sit near the centroid; they are not chosen to be worse solutions, and length or readability confounds are not reported separately. The program-simulation panel compares multi-model versus single-model corpora, not fingerprint selection from one pool. The coverage gap peaks at intermediate difficulty and intermediate sample budgets, then shrinks as both policies saturate. There is no experiment on non-GRPO RL, unverifiable rewards, or larger students.

Terms

Source

Related papers

All paper explainers