2-bit Bonsai-27B fits in 8GB VRAM but trails Qwen3.5-9B on Terminal-Bench

Creative-Regular6799 · reddit · 2026-07-21

A Reddit user benchmarked Ternary-Bonsai-27B at 2-bit and Bonsai-27B at 1-bit on Terminal-Bench 2.0 using an agent harness on an RTX 5070 Laptop with 8GB VRAM.

What they ran

Results

The key takeaway is that the 2-bit model does fit entirely on the GPU and tool calling stayed clean, but accuracy was worse than the smaller dense Qwen baseline. The 1-bit model was usable on simple prompts, but in an agent loop it produced a runaway 14,000+ token completion and never emitted a stop token, eventually exhausting the context window.

Original post →

More from coding & agent

coding & agent channel →