Community challenge pits MLX vs CUDA to speed up local Qwen 3.8 Flash on DGX Spark
gajesh · x · 2026-09-11
A community platform launched Qwen 3.8-Flash-Next with its first local-inference speed challenge, running the same model on two stacks with progress charted on one graph: a CUDA track using antirez's ds4 C/CUDA engine with Unsloth's GGUF quant on a single NVIDIA DGX Spark, and a Swift/Metal MLX track. "Let the optimizations begin."
More from coding & agent
- Tencent open-sources WeKnora, a RAG knowledge platform with 22k GitHub stars — abhishek__AI · 2026-09-11
- Tencent open-sources WeKnora, a RAG platform turning scattered docs into reasoning knowledge bases — abhishek__AI · 2026-09-11
- One Pokédex, four model setups: hands-on comparison of GPT-6 Astra and DeepSeek V4.1 Flash — kevinkern · 2026-09-11
- Agent Arena launches leaderboard grading models on millions of real agentic tasks — arena · 2026-09-11
- mcp-ecc: Open-Source MCP Server Unifying Email, Calendars and Contacts Across Providers — karljsamuel · 2026-09-11
- Codex Token usage surges post GPT-6 Astra; dev ships three Obsidian AI plugins — vista8 · 2026-09-11