Rust+Vulkan Inference Backend Goes Cross-Vendor: Pick Kernels by CPU ISA, Not GPU Vendor

PhysicsDisastrous462 · reddit · 2026-10-08

The author fixed a portability bug "hiding behind its own correctness" in their pure Rust + Vulkan Transformer backend: kernels tuned for Intel Gen9 had their reduction shapes baked into the supposedly portable path, breaking six fixtures on an AMD RDNA 3 handheld.

Root cause: the oracle is really the PyTorch CPU library on the host machine, and ATen dispatches vectorized kernels by instruction set at runtime — AVX2 hosts get 8-wide kernels, AVX-512 hosts 16-wide, changing the reduction shape. One ulp amplified through the gradient path pushed the model.embedtokens adjoint from 7.45e-9 to 3.22e-6, crossing the 2e-7 gate on six PEFT fixtures.

Fix: those two reductions are now selected by host CPU capability, probed once per process and cached, dispatching matching 8/16-lane modules; HIERARCHOSATENVECTORWIDTH=8|16 pins the shape for qualification. Both an Intel Skylake laptop and an AMD ROG Ally now pass 32/32 on LoRA, switching, and saved stages with the unchanged 2e-7 gate.

Verification: on Intel, the post-change report is bit-identical field-for-field across all 32 families; the Rust suite is 693 passed / 0 failed. The AMD side was re-qualified end to end with 3951 fingerprint inputs, zero changed or missing. Caveats: deterministic FP32 tiny-model correctness only, and AVX-512 is qualified solely on the AMD host.

Related event: Rust+Vulkan Inference Backend Fixes Cross-Vendor Portability Bug(2 posts)→

Original post →

More from Infra

Infra channel →