CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 24 upvotes

Selecting Diverse SFT Traces Improves Post-RL Generalization

QUESTION — How does reasoning-route diversity in SFT data affect model generalization after reinforcement learning?

This paper investigates route diversity—the variation in reasoning-step sequences in supervised fine-tuning (SFT) data—for preparing reasoning models for reinforcement learning (RL). The authors propose a lightweight, rule-based, CPU-only fingerprint to select for route diversity from a single pool. Experiments demonstrate that selecting diverse rather than similar routes enhances post-RL problem coverage across puzzles and mathematics, including on out-of-distribution tasks. Diagnostics reveal that diverse SFT yields both successful and failed attempts on more prompts, supplying group-relative RL with richer learning signals without incurring model-call overhead.

In synthetic experiments, route-diverse SFT improves OLMo3-7B's pass@8 by 16.9 points on environments held out from SFT.

In a single-model condition, where one model writes every candidate, diverse selection gains up to 6.2 points of mean pass@8 across 10 mathematics benchmarks.

On 3 open-source corpora, our CPU-only selector, without model calls, outperforms more expensive alternatives in every comparison of mean post-RL performance.

shizhuo2 · 27 Sept 2026 read the original ↗
↑