vLLM · builders · vllm ☆ FOLLOW 31 topic(s) over 14 day(s) — inference, architecture, evaluation. on X ↗
CARD NO. 2026-10-05 vLLM adds day-0 support for new open-weights models and scaling infrastructure inference · 3 voices
CARD NO. 2026-10-04 Cohere launches Embed 5 models and vLLM releases Semantic Router Decision 2.0 inference · 3 voices
CARD NO. 2026-09-28 Labs release new models and scientific agent benchmark leaderboards evaluation · 10 voicesvLLM adds day-one support for the multimodal MiMo-V2.6 model family inference · 4 voices
CARD NO. 2026-09-27 Grok 4.7 xHigh scores 58% on Artificial Analysis AA-Briefcase benchmark evaluation · 7 voicesApple releases a Qwen3.5-9B finetune converting documents into images architecture · 7 voicesvLLM adds day-one support for Xiaomi MiMo-V2.6 and DiffusionGemma-Jev inference · 4 voices
CARD NO. 2026-09-26 Grok 4.7 xHigh Reaches Competitive Scores on Benchmarks evaluation · 7 voicesApple and Developers Release New Open-Source Models inference · 7 voicesvLLM and SGLang add day-zero support for new Xiaomi and GLM models inference · 3 voices
CARD NO. 2026-09-25 DeepSeek V4.1 Flash launched alongside plans for a 2T model training architecture · 9 voicesQwen releases Qwen3.8-LiveTranslate for real-time interpretation multimodal · 7 voices
CARD NO. 2026-09-24 Release of Ternary Bonsai 2 27B and the Qwen3.8 real-time translation model inference · 8 voicesGrok 4.7 xHigh competes at the top alongside updates on Kimi K3 evaluation · 7 voicesDeepSeek advances large model development and outlines hardware training roadmap architecture · 6 voicesRunning local open-weight models on constrained hardware devices inference · 4 voicesOptimizing inference systems and accelerating execution for large language models inference · 3 voices
CARD NO. 2026-09-23 DeepSeek introduces V4.1-Flash and outlines future large model training plans inference · 8 voicesCommunity discusses Mixture of Experts models and hosting infrastructure architecture · 5 voicesAI community optimizes small models and expands local hardware support inference · 4 voicesXiaomi releases MiMo-V2.6 models with day-one vLLM support inference · 3 voices
CARD NO. 2026-09-22 Community anticipates Kimi K3 and Mixture of Experts models architecture · 7 voicesPrismML releases Ternary Bonsai 2 27B compressed from Qwen3.8 27B inference · 6 voicesXiaomi releases multimodal MiMo-V2.6 with day-0 vLLM support inference · 3 voices
CARD NO. 2026-09-21 Kimi prepares to release the Kimi K3.1 model architecture · 6 voicesPartnership brings TPU to vLLM alongside new Ternary model releases inference · 5 voices
CARD NO. 2026-09-20 vllm integrates TPU support and boosts model serving performance inference · 3 voices
CARD NO. 2026-09-19 Community discusses open-source AI models like GLM and Kimi ai · 9 voicesGoogle Cloud partners to bring TPU to vLLM alongside open-source releases system_design · 3 voices