Qwen 3.8 125B-A6B runs 15.3% faster on Mac via speculative decoding on mlx.fast
TheMoonMidas · x · 2026-09-11
- A new record on the mlx.fast optimization challenge: solver davidtai used speculative decoding to make Qwen 3.8 125B-A6B run 15.3% faster on Mac versus the official serial baseline.
- The official score is a single-stream composite (prefill gain^0.25 × decode gain^0.75), so decode carries 75% of the weight. The record run hits 36.3 TPS decode, 1585.8 TPS prefill, 1.4 tok/round, using a stock head with no extra training.
- Scores are computed and sealed by benchd; the CUDA leaderboard is still open for challengers.
More from Infra
- 1:26 continuous aerial AI video made entirely on a Mac with MiniMax H3 — cocktailpeanut · 2026-09-11
- KV cache gets QAT too: why this model beats others at fp4 KV cache — stochasticchasm · 2026-09-11
- Commentary: Anthropic loads shift to Google plus AWS slice, OpenAI doubles down on Azure — ericwdolan · 2026-09-11
- Nvidia claims Vera Rubin delivers 50X throughput per MW and 35X lower token cost vs Blackwell Ultra — Beth_Kindig · 2026-09-11
- Epoch AI: GPT long-context latency scales quadratically, matching price jumps — Jsevillamol · 2026-09-11
- RTX 3090 mini-bench: ninfer cuts TTFT from 3.4s to 29ms, prompt processing ~76x faster — milkipedia · 2026-09-11