Strata v0.1.40 adds Strix Halo support, multi-GPU batching and decode improvements
lxfater · x · 2026-10-06
Local inference engine Strata released v0.1.40 with experimental AMD Strix Halo / Ryzen AI Max (gfx1151) support: the HIP engine now runs on Radeon 8060S/8050S, unified memory is sized as one pool, UD-IQ4XS is auto-recommended, and Linux cold start is much faster. The release also brings multi-GPU batching improvements, better decode, Turing and Intel Arc gains, plus crash/NaN/security fixes. The project has 14.5k stars.
More from Infra
- BlackBerry's QNX hits record $80.3M revenue, up 27%, betting on AI-era safety platform — alysha_lobo · 2026-10-06
- Long-running benchmarks find Strata inference server failing full-build scenarios — julianharris · 2026-10-06
- AMD R9700 owner: ROCm on Windows cripples local i2v, Vulkan runs fine — vladomkd · 2026-10-06
- CutBCE: TPU Kernel Eliminates OOM in Large-Vocabulary Recommendation Training, 91.9% Faster — _reachsumit · 2026-10-06
- Domain specialists orchestrated by a general model: a local-LLM architecture pitch for 8-16GB GPUs — CyberExplore · 2026-10-06
- Run a 37GB Qwen MoE in a browser tab: LocalMind streams expert weights from disk, matches llama.cpp output — naklitechie · 2026-10-06