GGUF-based LoRA training fits Qwen3.5-35B-A3B into 16GB VRAM
woct0rdho · reddit · 2026-07-21
A GitHub project shows how to train Qwen3.5-35B-A3B LoRA models in 16 GB VRAM using GGUF as the base model format. The author claims GGUF is becoming a better base than bitsandbytes for low-VRAM training because it supports newer model types and quantization schemes, and reports that with APEX quant the model can fit at 13.3 GiB with fused dequant-matmul/MoE kernels. On Strix Halo, the setup reportedly reaches 6.5 s/it with batch size 1, 2048 token chunks, and LoRA rank 4, and the repo also spun out torch-ggml-ops for PyTorch bindings.
More from Infra
- NVIDIA starts rolling out 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-21
- Mustafa Suleyman says Microsoft is preparing for an OpenAI exit, while a new chip costs 30% less than GB200 — thoefler · 2026-07-21
- Speculative decoding boosts Qwen3.6-27B on one 5090, but slows crowded servers — luke_pacman · 2026-07-21
- NVIDIA says Blackwell Ultra hit 1,648 TFLOPs per GPU on DeepSeek-V3 671B training — NVIDIAAI · 2026-07-21
- A shopping app demo ties OpenTelemetry, Dynatrace and Port into agentic ops — Pavan_Belagatti · 2026-07-21
- Nvidia Rubin is coming, pointing to the next AI compute platform — ezyang · 2026-07-21