Rust + Vulkan training backend now parity-verifies 14 architectures and a full PEFT/LoRA workflow, on a handheld's iGPU

PhysicsDisastrous462 · reddit · 2026-09-30

An independent developer's native Rust + Vulkan Transformer training backend has grown from 7 to 14 parity-verified architectures in two weeks, including Falcon H1/H1R (parallel GQA/RoPE + per-layer Mamba2), DeepSeek V4, Phi-4 Multimodal, Kimi K2.5/K3, GPT-OSS, SmolLM3, Qwen models, Mistral 4, MiniMax M3, and Gemma 3/4.

Validation is strict: a pinned Hugging Face Transformers source tree serves as oracle, comparing forward logits, gradients, and two full AdamW steps with per-parameter checks at an absolute error ceiling of 2e-7 (worst observed: 1.19e-7, best: 2.6e-8). All local Vulkan validation ran on an ASUS ROG Ally Z1 Extreme's AMD RDNA 3 integrated GPU.

The bigger news: PEFT actually works now — HF-compatible LoRA adapter export, full modulestosave support (Linears, RMSNorm/LayerNorm, lmhead, embeddings) with multi-adapter switching and leak testing, bit-exact training resume, merge/unmerge, and a parameter-budget flag that auto-picks the largest LoRA rank within a base-model percentage. 32 architecture surfaces across 20 families pass all three PEFT stages with frozen-base drift exactly 0.0. The author notes COVID and an ear infection cost a few working days.

Related event: Solo Dev's Rust+Vulkan Training Backend Passes Validation on 14 Architectures(2 posts)→

Original post →

More from Infra

Infra channel →