Sharpening Tax in Post-Training
The paper investigates how reinforcement learning post-training affects language models, revealing that it pushes tasks toward extremes and trades solution coverage (pass@K) for single-shot accuracy (pass@1). The authors propose 'Sharpening Tax', a diagnostic metric quantifying the loss in test-time scalability. To mitigate this, they introduce posterior-tempered group sampling (PTGS), a Bayesian sampler adapting sampling temperature per prompt to estimated difficulty. Evaluated across 14 base and post-trained model pairs on three agentic benchmarks, PTGS incurs a lower tax and solves more tasks under repeated sampling.
Pre-trained LLMs with a light inference harness often surpass post-trained counterparts in solution coverage (pass@K) given sufficient test-time budget.
Sharpening Tax is a diagnostic metric that quantifies the loss in test-time scalability after post-training.
Posterior-tempered group sampling (PTGS) pays a smaller tax than fixed-temperature baselines, solving more tasks under repeated sampling.