CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 36 upvotes

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

QUESTION — How can Diffusion Large Language Models be accelerated while improving GPU memory efficiency during inference?

This work addresses memory I/O bottlenecks and parallel decoding limitations in Diffusion Large Language Models (dLLMs) that restrict their deployment. The authors introduce Flash-dLLM, a training-free inference acceleration framework featuring an I/O-aware fused KV-cache kernel and a unified KV-cache-driven draft-and-verify decoding strategy where the dLLM acts as both drafter and verifier. This design eliminates redundant memory movement and scales efficiently. Experiments on mathematical reasoning and code benchmarks demonstrate significant speedups over existing baselines such as Elastic-Cache.

It achieves 5.1times speedups over prior strongest baseline Elastic-Cache on GSM8K.

It achieves 11.0times speedups over prior strongest baseline Elastic-Cache on HumanEval.

mukul54 · 22 Sept 2026 read the original ↗
↑