Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
The paper investigates the impact of candidate-generation schedules on LLM test-time scaling, demonstrating that candidate count N alone is insufficient to describe system costs. The authors compare four schedules (1x8, 2x4, 4x2, and 8x1) while keeping N = 8 fixed. Results on A100 GPUs show that eight serial calls consume 4.64-4.86x as much gross GPU-device energy and exhibit 5.77-6.12x the P95 latency of a single batched call. Consequently, when candidates are independent and memory allows, fewer generation calls with larger batch sizes are more efficient.
On A100 GPUs, eight serial calls use 4.64-4.86x as much gross GPU-device energy as one batched call with eight candidates.
Eight serial calls have 5.77-6.12x the P95 latency of one batched call with eight candidates.
Increasing N from 1 to 8 improves accuracy by 8.4 percentage points for Phi-3-mini and 18.4 points for Qwen2.5-1.5B on 500 GSM8K prompts.