Local LLM field guide benchmarks 14 hardware configs for real-world token speed
NandoDF · x · 2026-10-11
A detailed October 2026 field guide maps 14 hardware configurations against 7 local models (plus 4 Claude Opus reference points), with measured decode speeds, memory, and cost data.
Key takeaways:
- Selection rule: pick the highest-scoring model whose decode range reaches 35 tok/s
- Qwen3.8-Flash-Next is called the sweet spot for capable local devices, running at 30+ tok/s across all major manufacturers' hardware
- Highlights: MacBook Pro M5 (32GB) runs Qwen3.6-35B-A3B 4-bit MLX at 51–60 tok/s; RTX 4090 runs Qwen3.8-27B Q4KS at 38–46 tok/s; RTX 5090 hits 59 tok/s; Ryzen AI Max+ 395 (128GB) manages 32–47 tok/s; DGX Spark uses NVFP4 with vLLM
A practical buying and tuning reference for anyone running local coding or tool-using agents.
More from Infra
- Compaction tosses 262GB of KV cache to keep 20KB, says Muennighoff — Muennighoff · 2026-10-11
- Qwen3.8-Flash-Next on a 5090 beats Claude Code on airbench: 11 min vs 14 min — dh7net · 2026-10-11
- Google's AI lag may stem from too little compute for short-term research, argues LeClerc — robleclerc · 2026-10-11
- Data center bottleneck shifts from demand to permission: Oracle's Project Jupiter stumbles — Beth_Kindig · 2026-10-11
- Energy capacity equals economic capacity: a16z chart fuels AI infra debate — _rockt · 2026-10-11
- Why buy $20k local machines for GLM 5.3 at 70 TPS? OpenRouter ran all night for $10 — TheZachMueller · 2026-10-11