Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches
The authors present Fathom, a key scan method where each query dynamically decides how many bits of each key channel to read from an offloaded 4-bit K cache stored as channel-major bit planes. By applying reverse water-filling over variance-weighted channel importance, queries optimize their bit budget. Running on Qwen3-8B at one million tokens, a decode step is 1.67x faster in GPU time than Double Sparsity, Loki, and SparQ r=32. Fathom reads 18% fewer bytes with lower attention error on six out of seven model and context settings while matching exact top-k decoding.
At one million tokens on Qwen3-8B a decode step is 1.67x faster in GPU time than with the 136-bit scans of Double Sparsity, Loki and SparQ r=32.
In the same GPU time as SparQ's 68-bit read (r=16) Fathom reads 18% fewer bytes with lower attention error on six of seven model and context settings.
On RULER-style tasks every per-token scan matches exact top-k decoding, and on real coding-agent sessions Fathom reaches the step agreement of the most accurate 136-bit scan at 92 bits.