Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
This paper introduces Parallel Power Tempering (PPT), an inference-time alternative to reinforcement learning post-training that enhances reasoning in LLMs through parallel power-sharpened sampling. By running multiple interacting replicas in parallel at different sharpening levels, PPT addresses the exploration-exploitation trade-off, allowing lower-power replicas to explore diverse trajectories and higher-power chains to exploit high-likelihood responses. The method mitigates a truncation bias found in prior samplers and investigates effective swap strategies. Extensive experiments show it substantially improves single-chain sampling and outperforms RL-post-trained models.
Parallel Power Tempering (PPT) instantiates power-sharpened LLM sampling via parallel tempering to resolve the exploration-exploitation trade-off at inference time.
PPT mitigates a truncation bias identified in prior power samplers under finite memory and compute budgets.
Extensive experiments show that PPT substantially improves single-chain power-sharpened sampling and outperforms RL-post-trained models.