ROCm Beats Vulkan on Strix Halo: Up to 47% Faster Decoding at 41.9 tok/s
QuixiAI · x · 2026-09-17
The Lucebox team reports pushing ROCm past Vulkan on AMD Strix Halo. Benchmarked same-day on the same machine with identical prompts: with the DSpark drafter, Lucebox (ROCm) decodes at 41.9 tok/s at 8K and 38.0 tok/s at 123K prompt tokens — 28-47% faster than llama.cpp/Vulkan v0.7.5, with 17-47% faster prefill. A notable win for local inference on AMD's AI Max hardware.
More from Infra
- Cloudflare Launches First Stealth Model Union Alpha, Blending Multiple LLMs Per Request — ritakozlov · 2026-09-17
- Qwen3.8-Flash-Next on SGLang NVFP4: 254K-Context TTFT Drops 35s→22s, Full Gauntlet Tested — FantasticNature7590 · 2026-09-17
- Privacy-focused LLM service Venice hits 250B daily tokens, up 2.5x in months — 0xAllen_ · 2026-09-17
- Baseten launches Hosted Tools, bringing server-side web search to open models — baseten · 2026-09-17
- Dev proposes predictive dynamic context caching for Claude Code — Sauers_ · 2026-09-17
- Starlink expands across Latin America: 10,000 antennas to connect 8,000 schools in Honduras alone — NicoVerderosa · 2026-09-17