Gemma 26B A4B Aces a C++ Coding Test Locally but Fumbles Tool Calls

HyperWinX · reddit · 2026-09-20

A Reddit user ran the Unsloth Gemma 26B A4B QAT Q4KXL GGUF locally with llama-server (MTP speculative decoding, 128K ctx, FP16 KV) and found it was the first model to one-shot their C++ HTTP server test. But tool calling in the Pi agent breaks down: it fails to submit plans and mixes up or forgets file-edit arguments, degrading badly past 30K context, with frequent repetition. Full llama-server flags included.

Original post →

More from coding & agent

coding & agent channel →