CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 5 upvotes

FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation

QUESTION — How can we eliminate semantic duplication in LLM text generation using temperature sampling without sacrificing accuracy?

LLM systems often use temperature sampling to aggregate completion samples, but this memoryless approach generates semantic duplicates. To fix this, the authors introduce FLEET, which integrates a memory mechanism into generation. FLEET tracks sparse trajectories through states where entropy exceeds a threshold, using them to compute per-token utility scores that adjust logits. Experiments show FLEET matches baseline accuracy with a 3x speedup, while increasing LiveCodeBench Pass@32 from 59.9% to 66.2% on complex coding tasks under equivalent budgets.

FLEET achieves the same accuracy as the repeated sampling baseline, with a 3x speedup.

LiveCodeBench Pass@32 increases from 59.9% to 66.2% under the same budget.

The approach is deterministic and uses a single calibration pass in the greedy-decoding configuration.

Alexiush · 23 Sept 2026 read the original ↗
↑