Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
This paper investigates test-time compute scaling for small reasoning models (sRMs), noting that self-refinement often consolidates probability onto already reachable solutions rather than discovering new ones. The authors identify execution bottlenecks and knowledge bottlenecks, introducing FlyBy—a selective querying framework where models reason, diagnose unresolved parts, and query stronger models only at knowledge bottlenecks. Using multi-depth supervised fine-tuning and cost-aware reinforcement learning to calibrate queries, the approach enables small models to match or exceed much larger models at a fraction of the serving cost.
On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%).
Scaling to FlyBy-8B further improves pass@8 to 51.81%.