27B Ternary Model Runs Within 10GB Memory
tcarambat · reddit · 2026-07-15
The post introduces PrismML's newly released Bonsai 27B: a Qwen-series model scaling the Ternary / BitNet approach to 27B parameters.
Key highlights include:
- Capable of running a 27B model on around 10GB of memory, with performance near fp16 precision at a 32K context length;
- Delivers significantly better results compared to similarly sized low-bit quantized versions, aiming to make large local models truly usable;
- They also demoed the model taking control of an entire computer via OpenComputer to conduct browser research and generate HTML reports.
The post also mentions unconfirmed capabilities and ecosystem details, such as dFlash, MTP support, 256K context, multimodal input, alongside ongoing llama.cpp and MLX branches. The author views this release as more significant than some higher-profile model updates, suggesting it marks a new milestone for local and edge-side open-source models.
Related event: Bonsai 27B: The first 27B model that runs on phones(15 posts)→
More from Infra
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11