PrismML introduces Ternary Bonsai 2 27B quantized model
PrismML has released Ternary Bonsai 2 27B based on Qwen3.8 27B, achieving a 9x reduction in size down to 5.9 GB while retaining 98.2% of its aggregate benchmark performance. Simultaneously, Qwen announced Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model built on an Interleave architecture.
6 independent accounts
43 posts
2 labs
112,497 interactions
qwen3kv cacheqwen3.8tokens per second