Fine-tuned Gemma 4 12B for 2.7x Better Tool Calling on 16GB VRAM

TheOneWhoWil · reddit · 2026-08-23

Unable to fit larger models on 16GB VRAM, the author fine-tuned Gemma 4 12B specifically for tool calling and CLI usage.

Results:

Assets: fp16 to Q4KM weights are available for use with llama.cpp or Ollama.

Original post →

More from coding & agent

coding & agent channel →