Solo Rust + Vulkan training backend now passes 14 architectures at 2e-7 parity, with full LoRA/PEFT lifecycle
PhysicsDisastrous462 · reddit · 2026-09-30
A solo developer updated their pure Rust + Vulkan Transformer training backend:
- 14 architectures now pass a strict parity gate (forward logits, gradients, two AdamW steps vs pinned HF Transformers, 2e-7 max error), including Falcon H1, DeepSeek V4, Kimi K2.5/K3, Qwen2.5/3.5/4-Exp, GPT-OSS, Mistral 4, MiniMax M2/M3, Gemma 3/4 and Phi-3/4; worst observed error 1.19e-7
- All local Vulkan validation ran on an ASUS ROG Ally Z1 Extreme's AMD RDNA 3 iGPU
- Full PEFT lifecycle works: HF-compatible LoRA export, modulestosave, bit-exact resume, merge/unmerge, multi-adapter isolation, automatic max-rank budgeting; 32 architecture surfaces pass all three PEFT gates with zero frozen-base drift
- The harness now fingerprints the pinned Transformers source alongside shaders to prevent silent reference drift
More from Infra
- Dumping GPUs and tokens can meaningfully speed up AI development — menhguin · 2026-09-30
- 9B Open-Weight Model Drops Agent Accuracy From 96% to 62.3% — TheZachMueller · 2026-09-30
- Jensen Huang: AI data centers add 10-20GW a year and roughly a million jobs — victor_explore · 2026-09-30
- First SGLang Summit set for Nov 12-13 in SF, with Intel CEO and Lilian Weng speaking — BanghuaZ · 2026-09-30
- Interview with Richard Ho, leader of OpenAI's in-house chip project — bookwormengr · 2026-09-30
- Tenstorrent opens bio-model training: OpenFold3 on Blackhole Galaxy nears DGX H200 at quarter the cost — MoAlQuraishi · 2026-09-30