Bonsai 2 27B quantized beats Gemma 4 12B and Qwen 3.5 9B in 7GB 3D generation test
Fun-Meaning-6474 · reddit · 2026-09-19
A Reddit user benchmarked PrismML's heavily quantized Bonsai 2 27B (PQ20, 7.2GB) against Gemma 4 12B (Q80) and Qwen 3.5 9B (Q6K) on a single RTX 5090, giving all three the same 770-word prompt to generate a voxel Japanese pagoda scene in three.js as a single HTML file, with 262K context.
- Bonsai 2 27B: 106,396 output tokens, 17.5 min, 101 tok/s, 90% of budget on thinking, clearly the best detail
- Gemma 4 12B: 9,159 tokens, 95s, 96 tok/s, decent but washed out
- Qwen 3.5 9B: 12,105 tokens, 70s, 168 tok/s, output largely broken
The author concludes Bonsai 2 offers unmatched intelligence per GB and can run decently on an RTX 3060 — a big win for local AI. Note it requires PrismML's llama.cpp fork to load; stock llama.cpp fails. The author discloses being a co-founder of atomic.chat, which uses PrismML's backend.
More from Infra
- HEIF Heist: one C image parser bug chain leads to RCE in OpenAI, Meta, GitHub — ccerrato147 · 2026-09-19
- Turbo-dLLM Open-Sources CSBP, Speeding Diffusion LLM Training Up to 7.59x at 1M Context — Azaliamirh · 2026-09-19
- Luminal runs large-scale FLUX.2 diffusion on AMD MI300X to cut cost per image — ycombinator · 2026-09-19
- Virginia governor creates AI task force and moves to restrain data centers — The Verge AI · 2026-09-19
- Google Cloud adds native transactional queues to Spanner for in-database async work — rseroter · 2026-09-19
- AI Infra Summit: Penguin Solutions and Astera Labs Bet Big on CXL Memory Expansion — BenBajarin · 2026-09-19