CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 54 upvotes

Disaggregated Quantization: Specializing LLM Prefill and Decode

QUESTION — How can quantization formats be specialized for LLM prefill and decode phases to improve inference efficiency?

The paper proposes "disaggregated quantization" (DQ) to specialize computation formats, weights, and storage placement for LLM prefill and decode phases. By removing activation quantization on decode and training compute-native prefill weights, the approach accelerates prompt processing while preserving accuracy. An offloaded disaggregated prefill (ODP) system streams weights from SSD to fit single devices, delivering substantial speedups in llama.cpp and vLLM across large model scales.

With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint.

On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp.

We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.

BlackSamorez · 22 Sept 2026 read the original ↗
↑