SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops
The study introduces SiliconBench to evaluate nine Apple Silicon serving engines across speed, memory, and fidelity using chat and agent workloads on Qwen3, Qwen3.5, and Gemma 4 models. The findings reveal that explicit memory budgets do not guarantee memory headroom, and scheduling prompt processing alongside ongoing generation is crucial for maintaining low first-token latency. Furthermore, the evaluation shows that tensor parallelism over Thunderbolt RDMA scales effectively while pipeline parallelism over TCP degrades under multi-machine configurations.
vllm-metal alone more than doubles throughput on both workloads from concurrency 1 to 16 on Qwen3-0.6B.
Only three stacks satisfy the completion, fidelity, and model-coverage gates.
Tensor parallelism over Thunderbolt RDMA scales while pipeline parallelism over TCP regresses.