Ternary-Bonsai-2-27B 2-bit MLX quantized model trends on Hugging Face
prism-ml · hf · 2026-09-18
prism-ml's Ternary-Bonsai-2-27B MLX 2-bit quantized model is trending on Hugging Face. It targets on-device text generation with ternary quantization, hybrid attention (prismhadamardqwen35), and CUDA/Metal support, squeezing a 27B-class model into 2-bit for local inference.
Related event: PrismML releases Ternary Bonsai 2 27B: 5.95GB, retains 98.2% performance(11 posts)→
More from Infra
- 0.05% sampling to validate cache hits: developer marvels at compute saved across the system — DanielLockyer · 2026-09-18
- TRL Adds Async GRPO with LoRA Weight Sync over HF Buckets, Cutting Training from 3.5h to 53min — _lewtun · 2026-09-18
- Payments firms race to own AI inference: Stripe taps OpenRouter, Ramp enters the chain — xkonjin · 2026-09-18
- Qdrant wraps 4+ hour Vector Space Stream on vector search — recording now live — qdrant_engine · 2026-09-18
- Investor: QNX's microkernel is an undervalued moat for the agentic AI era — pdamodaran · 2026-09-18
- Inception CEO Stefano Ermon argues diffusion will beat autoregressive models on inference efficiency — No Priors · 2026-09-18