FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation
LLM systems often use temperature sampling to aggregate completion samples, but this memoryless approach generates semantic duplicates. To fix this, the authors introduce FLEET, which integrates a memory mechanism into generation. FLEET tracks sparse trajectories through states where entropy exceeds a threshold, using them to compute per-token utility scores that adjust logits. Experiments show FLEET matches baseline accuracy with a 3x speedup, while increasing LiveCodeBench Pass@32 from 59.9% to 66.2% on complex coding tasks under equivalent budgets.
FLEET achieves the same accuracy as the repeated sampling baseline, with a 3x speedup.
LiveCodeBench Pass@32 increases from 59.9% to 66.2% under the same budget.
The approach is deterministic and uses a single calibration pass in the greedy-decoding configuration.