Ternary-Bonsai-2-27B, a 2-bit ternary model for on-device inference, trends on Hugging Face
prism-ml · hf · 2026-09-18
prism-ml's Ternary-Bonsai-2-27B is trending on Hugging Face as an open text-generation model.
- Released in GGUF format, runnable directly in llama.cpp with CUDA and Metal support
- Uses ternary (2-bit) weights aimed at local/on-device deployment
- Features a hybrid-attention architecture for end-side efficiency
Related event: PrismML Releases Ternary Bonsai 2 27B: 5.95GB Model Runs in Browser(7 posts)→
More from Infra
- 600 tok/s single-request on Qwen 35B with Ninfer on an RTX Pro 6000 — CharlesStross · 2026-09-18
- Cadence sees India's EDA market doubling to $7.82B by 2031 — bookwormengr · 2026-09-18
- Google Open-Sources Agent Substrate on GKE: 10x Density, 1,000+ Dormant Agents per Host — blaizedsouza · 2026-09-18
- Redditor crams six V100 GPUs into a standard full-tower case for local LLM inference — Odd_Caterpillar_2994 · 2026-09-18
- Crusoe raises $3.9B at $30.9B valuation to build data centers and modular AI factories — TechCrunch AI · 2026-09-18
- A 2.5-hour first-principles primer on the semiconductor supply chain worth your time — blaizedsouza · 2026-09-18