27B Ternary Model Runs Within 10GB Memory
tcarambat · reddit · 2026-07-15
The post introduces PrismML's newly released Bonsai 27B: a Qwen-series model scaling the Ternary / BitNet approach to 27B parameters.
Key highlights include:
- Capable of running a 27B model on around 10GB of memory, with performance near fp16 precision at a 32K context length;
- Delivers significantly better results compared to similarly sized low-bit quantized versions, aiming to make large local models truly usable;
- They also demoed the model taking control of an entire computer via OpenComputer to conduct browser research and generate HTML reports.
The post also mentions unconfirmed capabilities and ecosystem details, such as dFlash, MTP support, 256K context, multimodal input, alongside ongoing llama.cpp and MLX branches. The author views this release as more significant than some higher-profile model updates, suggesting it marks a new milestone for local and edge-side open-source models.
Related event: Bonsai 27B: The first 27B model that runs on phones(15 posts)→
More from Infra
- Engram shows how agent memory can keep, rewrite, or delete facts asynchronously — philipvollet · 2026-07-21
- Lightning AI’s LitLogger captures training metrics, artifacts, commands, and environment data — LightningAI · 2026-07-21
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- Microsoft expands Mistral models across Azure, Foundry, Copilot Studio and Azure Local — arthurmensch · 2026-07-21
- SmolVM is pitched as a lighter in-house sandbox for agent runtimes — aniketmaurya · 2026-07-21
- Three-part PyTorch profiling series explains torch.profiler for accelerator debugging — RisingSayak · 2026-07-21