Low-VRAM LoRA training now fits DeepSeek-V4-Flash into 90 GiB VRAM
woct0rdho · reddit · 2026-07-29
- The author reports progress on low-VRAM LoRA training over GGUF base models, and says DeepSeek-V4-Flash (284B-A13B) can now be trained in 90 GiB VRAM with no CPU offloading.
- On Strix Halo, the setup runs at 19 s/it.
- The post says the difficult parts—sliding attention, CSA, and HCA—now use vibe-coded Triton kernels that are reportedly faster than other implementations the author has seen.
- Beyond training, the author hopes the PyTorch/GGUF integration work will help enable more research into non-training model surgery, citing Heretic and noting that mHC ablation remains an unsolved task.
More from Infra
- Podcast spotlights the future of vector databases in the AI stack — CShorten30 · 2026-07-29
- Qdrant, Future AGI and AWS set a talk on self-improving agent retrieval loops — qdrant_engine · 2026-07-29
- Google opens early-access Gemini distillation service for smaller, cheaper models — ccerrato147 · 2026-07-29
- A 27B model reaches 24 TPS with on-the-fly 3-bit dequantization on an A6000 — cephaloform · 2026-07-29
- OpenRouter’s weighted average token price fell sharply this year — maferase · 2026-07-29
- Decentralized Network Trains 16B Model Across 3 Continents Using RTX 4090s — bittingthembits · 2026-07-29