Rust + Vulkan training backend now parity-verifies 14 architectures and a full PEFT/LoRA workflow, on a handheld's iGPU
PhysicsDisastrous462 · reddit · 2026-09-30
An independent developer's native Rust + Vulkan Transformer training backend has grown from 7 to 14 parity-verified architectures in two weeks, including Falcon H1/H1R (parallel GQA/RoPE + per-layer Mamba2), DeepSeek V4, Phi-4 Multimodal, Kimi K2.5/K3, GPT-OSS, SmolLM3, Qwen models, Mistral 4, MiniMax M3, and Gemma 3/4.
Validation is strict: a pinned Hugging Face Transformers source tree serves as oracle, comparing forward logits, gradients, and two full AdamW steps with per-parameter checks at an absolute error ceiling of 2e-7 (worst observed: 1.19e-7, best: 2.6e-8). All local Vulkan validation ran on an ASUS ROG Ally Z1 Extreme's AMD RDNA 3 integrated GPU.
The bigger news: PEFT actually works now — HF-compatible LoRA adapter export, full modulestosave support (Linears, RMSNorm/LayerNorm, lmhead, embeddings) with multi-adapter switching and leak testing, bit-exact training resume, merge/unmerge, and a parameter-budget flag that auto-picks the largest LoRA rank within a base-model percentage. 32 architecture surfaces across 20 families pass all three PEFT stages with frozen-base drift exactly 0.0. The author notes COVID and an ear infection cost a few working days.
More from Infra
- Dual R9700 local inference: how much does PCIe Gen4 actually cost vs Gen5? — IngwiePhoenix · 2026-09-30
- SwitchSD, speculative decoding that switches between neural drafting and context copying, accepted to NeurIPS — arankomatsuzaki · 2026-09-30
- Rebellions' 2048 TFLOPS NPU lands major Japanese AI datacenter deal with ai& — DavidBennett__ · 2026-09-30
- Solo Rust + Vulkan training backend now passes 14 architectures at 2e-7 parity, with full LoRA/PEFT lifecycle — PhysicsDisastrous462 · 2026-09-30
- Micron guided $50B quarterly revenue; independent model says $56.2B if price hikes pass through — tengyanAI · 2026-09-30
- DeepSeek ships open-source TileLang tools for Huawei Ascend chips, countering CUDA — The Decoder · 2026-09-30