Running local open-weight models on constrained hardware devices
Developers deployed a 27-billion-parameter ternary model at 1.72 bits per weight on a single graphics card with 12GB of VRAM. Additionally, updates brought llama.cpp's Metal kernels to transformers and added SGLang support for image generation models.
4 independent accounts
20 posts
3 labs
11,277 interactions
llama.cppllamaquantizationunslothgemma