Finetuning away GQA: Qwen 3.8 27B experiments and call for collaboration
Signature97 · reddit · 2026-09-02
The author experimented with finetuning Qwen 3.8 27B to replace GQA layers with KDA layers, aiming to reduce KV cache overhead. Inspired by Arcee's DistilKit, the strategy was tested on 26.2K tokens, showing poor performance—specifically a significant drop on GSM8K—mirroring challenges noted by Arcee. Due to limited personal compute, the author is calling for community collaboration to pool resources for larger-scale finetuning.
More from Infra
- Anthropic Cuts Cache Read Prices by 75%, Closing Gap with DeepSeek — Teknium · 2026-09-02
- Trading sector too small to justify massive AI capex, analyst argues — iaindunning · 2026-09-02
- Speed Wars: GPT 5.6 Hits 750 Tokens/s on Cerebras — DeepLearningAI · 2026-09-02
- aimake: Incremental Build System for AI/ML Pipelines — Miserable_Extent8845 · 2026-09-02
- Fable 5.1 Available on Hermes Agent and OpenRouter — Scobleizer · 2026-09-02
- ComfyUI nodes trigger RTX 5070Ti system crashes — Pitiful-Indication95 · 2026-09-02