Ternary 2-bit Bonsai-2-27B GGUF lands on Hugging Face trending
dealignai · hf · 2026-09-20
Bonsai-2-27B-Ternary-CRACK-GGUF, a ternary 2-bit quantized 27B model from PrismML, is trending on Hugging Face. It pairs ternary quantization with a hybrid-attention architecture and targets on-device inference via llama.cpp, CUDA, Metal and ONNX, aiming to run 27B-scale models cheaply on consumer hardware.
More from Infra
- SpaceX's 100,000-satellite Starlink V3 constellation accepted by the FCC — ns123abc · 2026-09-20
- jevcache memoizes model decisions by (model, schema, state) — repeat calls cost $0 and return in ~0ms — JiliJeanlouis · 2026-09-20
- Key takeaways from the 2026 AI Infra Summit with 9,000+ attendees — karlfreund · 2026-09-20
- Brain Runs on 20 Watts: Can Neuromorphic Computing Make AI Less Power-Hungry? — burny_tech · 2026-09-20
- Pedro Domingos: ASML Is a Single Point of Failure for the AI Supply Chain — pmddomingos · 2026-09-20
- LexiPanel: open-source AIO web panel for llama.cpp inference servers — W61k3r · 2026-09-20