CONSONANCE.for your information
Thứ Hai, 5 tháng 10, 2026frenvi

Đáng đọc kỹ

01 — inference 54 upvote

Disaggregated Quantization: Specializing LLM Prefill and Decode

CÂU HỎI — Làm thế nào để chuyên biệt hóa định dạng lượng tử hóa cho từng pha prefill và decode của mô hình ngôn ngữ lớn nhằm cải thiện hiệu năng inference?

Bài báo đề xuất phương pháp "disaggregated quantization" (DQ) nhằm chuyên biệt hóa định dạng tính toán, trọng số và vị trí lưu trữ cho hai pha prefill và decode của LLM. Bằng cách loại bỏ lượng tử hóa activation trong quá trình decode và huấn luyện trọng số prefill chuyên biệt cho tính toán, phương pháp này tăng tốc độ xử lý prompt mà vẫn duy trì hoặc vượt độ chính xác. Hệ thống offloaded disaggregated prefill (ODP) truyền trọng số từ SSD để xử lý trên một thiết bị đơn lẻ, mang lại tốc độ cải thiện đáng kể trong llama.cpp và vLLM.

With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint.

On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp.

We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.

BlackSamorez · 22 thg 9, 2026 đọc bản gốc ↗
↑