PrismML’s Bonsai 27B ternary runs on an RX 9070 XT and survives real tool calls

blakok14 · reddit · 2026-07-29

A Reddit user tested PrismML’s Bonsai 27B (ternary) on an RX 9070 XT 16GB via llama.cpp, focusing on real tool-use rather than benchmarks. In a structured code-edit workflow, the model’s reasoning held up better than expected for a heavily compressed model, and it could call tools successfully.

The downside was reliability: it still produced enough syntax errors that the author wouldn’t trust it unsupervised for serious agentic work. Even so, the combination of 1.7 bits per weight compression, 262K context, and functioning tool-calling was described as a genuine milestone. The author also asked for raw throughput numbers on AMD/RDNA hardware.

Original post →

More from coding & agent

coding & agent channel →