Finetuning away GQA: Qwen 3.8 27B experiments and call for collaboration

Signature97 · reddit · 2026-09-02

The author experimented with finetuning Qwen 3.8 27B to replace GQA layers with KDA layers, aiming to reduce KV cache overhead. Inspired by Arcee's DistilKit, the strategy was tested on 26.2K tokens, showing poor performance—specifically a significant drop on GSM8K—mirroring challenges noted by Arcee. Due to limited personal compute, the author is calling for community collaboration to pool resources for larger-scale finetuning.

Original post →

More from Infra

Infra channel →