LoRA over GGUF: fine-tune Qwen3.8-Flash-Next in 40 GiB VRAM without CPU offloading
woct0rdho · reddit · 2026-10-09
The author updates their LoRA over GGUF series: Qwen3.8-Flash-Next (125B-A6B + 51B engram) can now be trained in 40 GiB VRAM with no CPU offloading, with engram kept on disk without slowing training.
On Strix Halo, context chunk size 2048 runs at 9.5 s/it (200 token/s). There's still headroom versus >1600 token/s prefill, and LoRA training typically costs 2-3x prefill work. Modern GGUF initial support merged in transformers 5.18, with more to come.
Spoiler: GGTensile in torch-ggml-ops does Tensile-like asm-level optimization on MMQ kernels, already beating HIP in many cases.
More from Infra
- Bittensor-based GPU cloud Lium buys back and burns nearly $2.7M of SN51 tokens in six months — markjeffrey · 2026-10-09
- Google open-sources ML Drift: one GPU engine for GLES/OpenCL/Metal/WebGPU, cutting Shorts frame latency 40% — lmoroney · 2026-10-09
- Inference platform doubles providers since launch as competition pushes prices down — AccBalanced · 2026-10-09
- Liquid Inference doubles provider network in one day, undercuts OpenRouter on 30+ models — AccBalanced · 2026-10-09
- SparseEngine: sparse-first inference engine delivers 10x throughput with KV eviction — Jitai Hao · 2026-10-09
- Data centers use ~1 km³ of water a year vs hundreds for households — the clash is location, not volume — AryHHAry · 2026-10-09