Selecting Diverse SFT Traces Improves Post-RL Generalization
This paper investigates route diversity—the variation in reasoning-step sequences in supervised fine-tuning (SFT) data—for preparing reasoning models for reinforcement learning (RL). The authors propose a lightweight, rule-based, CPU-only fingerprint to select for route diversity from a single pool. Experiments demonstrate that selecting diverse rather than similar routes enhances post-RL problem coverage across puzzles and mathematics, including on out-of-distribution tasks. Diagnostics reveal that diverse SFT yields both successful and failed attempts on more prompts, supplying group-relative RL with richer learning signals without incurring model-call overhead.
In synthetic experiments, route-diverse SFT improves OLMo3-7B's pass@8 by 16.9 points on environments held out from SFT.
In a single-model condition, where one model writes every candidate, diverse selection gains up to 6.2 points of mean pass@8 across 10 mathematics benchmarks.
On 3 open-source corpora, our CPU-only selector, without model calls, outperforms more expensive alternatives in every comparison of mean post-RL performance.