Qwen 3.8 Flash Next Optimization Challenge: MLX and CUDA Both Gain Over 55%

Alibaba's Qwen launched Qwen 3.8-Flash-Next on Spark and MLX communities with a joint local inference optimization challenge, pitting MLX against CUDA on the same model. Both tracks have already achieved over 55% speedups.

2026-09-11 ~ 2026-09-11 · 2 related posts