MLX vs CUDA: Qwen 3.8 Flash Next optimization duel delivers 55%+ speedups on both sides
gajesh · x · 2026-09-11
- Alibaba's Qwen team launched Qwen 3.8-Flash-Next on both the Spark and MLX communities as an optimization challenge, tracking both communities' progress on the same graph for the first time.
- The task: make Qwen 3.8 Flash Next run faster on two concrete setups — a C/CUDA engine ported from antirez's ds4 (with Unsloth's GGUF weights on a single NVIDIA DGX Spark) versus a Swift/Metal implementation.
- Both challenges have already delivered over 55% speed improvement, with CUDA currently holding a slight edge. The author frames it as friendly competition between two highly overlapping communities pushing the same model as far as possible.
Related event: Qwen 3.8 Flash Next Optimization Challenge: MLX and CUDA Both Gain Over 55%(2 posts)→
More from Infra
- One cheeseburger emits as much CO2 as 63,000 Gemini text prompts, math shows — recallingmemories · 2026-09-11
- spcx reportedly signed another mega compute deal a week ago, $13B ARR per CFO — rwang07 · 2026-09-11
- PyTorch Lightning checkpointing runs up to 95% faster on Google Cloud — LightningAI · 2026-09-11
- LuxTTS: open-source voice cloning at 48kHz, 150x realtime, under 1GB VRAM — tom_doerr · 2026-09-11
- Nearly 10% of exposed LiteLLM gateways accept default admin key 'sk-1234' — Thionne_WTZ · 2026-09-11
- K2 Horizon's frozen-model LoRA gives 3x faster inference, trained on 20T tokens — rohanpaul_ai · 2026-09-11