CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 37 upvotes

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

QUESTION — How does the candidate-generation strategy impact the energy consumption and performance of LLM test-time scaling systems?

The paper investigates the impact of candidate-generation schedules on LLM test-time scaling, demonstrating that candidate count N alone is insufficient to describe system costs. The authors compare four schedules (1x8, 2x4, 4x2, and 8x1) while keeping N = 8 fixed. Results on A100 GPUs show that eight serial calls consume 4.64-4.86x as much gross GPU-device energy and exhibit 5.77-6.12x the P95 latency of a single batched call. Consequently, when candidates are independent and memory allows, fewer generation calls with larger batch sizes are more efficient.

On A100 GPUs, eight serial calls use 4.64-4.86x as much gross GPU-device energy as one batched call with eight candidates.

Eight serial calls have 5.77-6.12x the P95 latency of one batched call with eight candidates.

Increasing N from 1 to 8 improves accuracy by 8.4 percentage points for Phi-3-mini and 18.4 points for Qwen2.5-1.5B on 500 GSM8K prompts.

iammobina · 16 Sept 2026 read the original ↗
↑