LoRA over GGUF: fine-tune Qwen3.8-Flash-Next in 40 GiB VRAM without CPU offloading

woct0rdho · reddit · 2026-10-09

The author updates their LoRA over GGUF series: Qwen3.8-Flash-Next (125B-A6B + 51B engram) can now be trained in 40 GiB VRAM with no CPU offloading, with engram kept on disk without slowing training.

On Strix Halo, context chunk size 2048 runs at 9.5 s/it (200 token/s). There's still headroom versus >1600 token/s prefill, and LoRA training typically costs 2-3x prefill work. Modern GGUF initial support merged in transformers 5.18, with more to come.

Spoiler: GGTensile in torch-ggml-ops does Tensile-like asm-level optimization on MMQ kernels, already beating HIP in many cases.

Original post →

More from Infra

Infra channel →