Local MoE Benchmark: NVIDIA Lightning Outruns Qwen by 2.5x
parepeg · reddit · 2026-08-12
Based on Artificial Analysis data and local llama.cpp testing, the author compares two MoE models: NVIDIA Lightning (30B) and Qwen3.6 (35B).
- Speed vs. Intelligence: NVIDIA Lightning is about 2.5x faster per task, ideal for quick responses without complex reasoning. Qwen is more intelligent but tends to think for a set amount of time regardless of complexity.
- Local Test Data (AMD Strix Halo 395, 20k tokens):
- Lightning 30B: Prefill 1059 tok/s, Decode 53.48 tok/s.
- Qwen3.6 35B (MTP enabled): Prefill 988.78 tok/s, Decode 50.98 tok/s.
More from Infra
- Linux Cloud GPU Blocked from Video Upscaling: NVIDIA RTX VSR is Windows-Only — emacrema · 2026-08-12
- GPU Demand Surges: H100 Rental Prices Jump 40% in Six Months — jessi_cata · 2026-08-12
- AI Data Center Firm DayOne Confidentially Files for US IPO — davidyin44 · 2026-08-12
- Anthropic Reportedly Inks $9.1B Compute Deal; Microsoft to Massively Boost AI Chip Production — 创业邦 · 2026-08-12
- AI Chip Yield Anxiety: Why Wafer-Level Testing is Becoming Critical — demian_ai · 2026-08-12
- RTX 4080 hits OOM running Minimax Ref2V: How to generate long videos locally? — witcherknight · 2026-08-12