PrismML’s Bonsai 27B ternary runs on an RX 9070 XT and survives real tool calls
blakok14 · reddit · 2026-07-29
A Reddit user tested PrismML’s Bonsai 27B (ternary) on an RX 9070 XT 16GB via llama.cpp, focusing on real tool-use rather than benchmarks. In a structured code-edit workflow, the model’s reasoning held up better than expected for a heavily compressed model, and it could call tools successfully.
The downside was reliability: it still produced enough syntax errors that the author wouldn’t trust it unsupervised for serious agentic work. Even so, the combination of 1.7 bits per weight compression, 262K context, and functioning tool-calling was described as a genuine milestone. The author also asked for raw throughput numbers on AMD/RDNA hardware.
More from coding & agent
- Kimi K3 reportedly works well in Kimi Code and Claude Code via the Responses API — zainhas · 2026-07-29
- Figma and Sentry MCP servers push design data and live errors into coding agents — heypearlai · 2026-07-29
- GitHub MCP Server and Context7 emerge as core tools for coding agents — heypearlai · 2026-07-29
- Cutting half your MCP servers may make your agent smarter overnight — heypearlai · 2026-07-29
- Alexey Grigorev’s AI dev workshop covers specs, tests, Docker, and CI/CD — Al_Grigor · 2026-07-29
- AI coding agents may be killing the developer flow state — bendee983 · 2026-07-29