CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 18 upvotes

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

QUESTION — How can large Mixture-of-Experts models be served from an SSD on consumer hardware without hitting the weight memory bottleneck?

The authors introduce Edge0, a streaming MoE inference engine that solves memory bottlenecks when offloading a 35B model to an SSD by utilizing a prerouter. This prerouter predicts the next layer's routing one token ahead, enabling SSD reads to start early enough to hide behind compute. Additionally, an unmerged recovery LoRA trained on the student path compensates for quality loss from int4 quantization. Results show that on a single 24GB machine, Edge0 serves a 35B MoE at 20tok/s within 3GiB of peak active memory.

Edge0 uses a prerouter to predict next-layer routing one token ahead to overcome SSD offloading latency.

An unmerged recovery LoRA compensates for quality lost to int4 quantization and routing replacement.

On a single 24GB machine, Edge0 serves a 35B MoE at 20tok/s inside 3GiB of peak active memory.

bupalinyu · 16 Sept 2026 read the original ↗
↑