CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 6 upvotes

SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops

QUESTION — How do various LLM serving engines perform across speed, memory, and fidelity on unified-memory desktops?

The study introduces SiliconBench to evaluate nine Apple Silicon serving engines across speed, memory, and fidelity using chat and agent workloads on Qwen3, Qwen3.5, and Gemma 4 models. The findings reveal that explicit memory budgets do not guarantee memory headroom, and scheduling prompt processing alongside ongoing generation is crucial for maintaining low first-token latency. Furthermore, the evaluation shows that tensor parallelism over Thunderbolt RDMA scales effectively while pipeline parallelism over TCP degrades under multi-machine configurations.

vllm-metal alone more than doubles throughput on both workloads from concurrency 1 to 16 on Qwen3-0.6B.

Only three stacks satisfy the completion, fidelity, and model-coverage gates.

Tensor parallelism over Thunderbolt RDMA scales while pipeline parallelism over TCP regresses.

windchimeran · 12 Sept 2026 read the original ↗
↑