Community challenge pits MLX vs CUDA to speed up local Qwen 3.8 Flash on DGX Spark

gajesh · x · 2026-09-11

A community platform launched Qwen 3.8-Flash-Next with its first local-inference speed challenge, running the same model on two stacks with progress charted on one graph: a CUDA track using antirez's ds4 C/CUDA engine with Unsloth's GGUF quant on a single NVIDIA DGX Spark, and a Swift/Metal MLX track. "Let the optimizations begin."

Original post →

More from coding & agent

coding & agent channel →