Optimizing Qwen3.8-27B on MI300X boosts speed by 59%
cheptsov · reddit · 2026-08-21
dstack demonstrates optimizing Qwen3.8-27B inference on a single AMD MI300X using their open-source toolkit. By linking optimization sessions and applying source-level patches to SGLang's AITER attention backend, they increased inference speed from 311 to 495 tok/s (+59%). The optimized setup supports a 1M context with p50 TTFT under 1.5s and handles four concurrent users (10k in / 1.5k out). The result is a portable preset deployable on any AMD cloud, Kubernetes cluster, or bare-metal fleet.
More from Infra
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24