Bonsai 2 27B quantized beats Gemma 4 12B and Qwen 3.5 9B in 7GB 3D generation test

Fun-Meaning-6474 · reddit · 2026-09-19

A Reddit user benchmarked PrismML's heavily quantized Bonsai 2 27B (PQ20, 7.2GB) against Gemma 4 12B (Q80) and Qwen 3.5 9B (Q6K) on a single RTX 5090, giving all three the same 770-word prompt to generate a voxel Japanese pagoda scene in three.js as a single HTML file, with 262K context.

The author concludes Bonsai 2 offers unmatched intelligence per GB and can run decently on an RTX 3060 — a big win for local AI. Note it requires PrismML's llama.cpp fork to load; stock llama.cpp fails. The author discloses being a co-founder of atomic.chat, which uses PrismML's backend.

Original post →

More from Infra

Infra channel →