Gemma 3 27B QAT Reported to Duplicate Tool Calls in Local Setups
Mrinohk · reddit · 2026-08-04
A developer reported that running the Gemma 3 27B QAT model locally with tool-calling enabled leads to frequent duplicate tool invocations. In a harness preserving reasoning context, the model gets confused after receiving the initial tool response and triggers the call again. The author notes that Qwen2.5 32B does not exhibit this behavior, suspecting it might be related to llama.cpp's tool schema handling or specific model quirks.
More from coding & agent
- Liquid AI Launches LFM2.5-2.6B: On-Device Agentic Model Outperforming Larger Counterparts — maximelabonne · 2026-08-04
- Balancing Privacy and Power: Building a Hybrid Local-Cloud LLM Workflow — KhuyenTran16 · 2026-08-04
- Building From Anywhere: A Mac Mini + Terminal Setup for Parallel AI Agents — EXM7777 · 2026-08-04
- Cognition President Shares Insights on Building Proactive AI Agents — LangChain · 2026-08-04
- Cloudflare Proposes Shift to Agent Development Lifecycle with Full-Stack Tools — ritakozlov · 2026-08-04
- Tutorial: AI-Native Developer Workflow for Using AI Tools Without Losing Control — Al_Grigor · 2026-08-04