AMD iGPU LLM Benchmark: MoE Architectures Crush Dense Models on Edge Inference
tabletuser_blogspot · reddit · 2026-08-01
The poster benchmarked various mainstream MoE and Dense models using llama.cpp (Vulkan backend) on a Mini PC with an AMD Ryzen 7 6800H iGPU (Radeon 680M, 1GB assigned VRAM, 64GB RAM).
Key Findings:
- MoE Dominance: Mixture-of-Experts (MoE) models significantly outperform traditional Dense models on integrated GPU systems.
- Performance Highlights: gpt-oss 20B (Q6K) achieved 353 pp/t/s and 16.85 tg/t/s; gemma4 26B.A4B (Q40) reached 312 pp/t/s and 18.35 tg/t/s.
- Slow Dense Models: Dense models like qwen35 27B and gemma4 31B generally outputted less than 3 t/s (tg128).
- The author notes a lack of 70B-level MoE models and failed to load Qwen3-Coder-Next-MXFP4MOE.
More from Infra
- Lost Tacit Knowledge Threatens Rapid Nuclear Buildout for AI — gabriel1 · 2026-08-01
- Qualcomm Acquires Modular to Tackle AI Software Bottlenecks — clattner_llvm · 2026-08-01
- Elon Musk Declares >99.99% of AI Computing Will Eventually Move to Space — Polymarket · 2026-08-01
- CoreWeave Launches Sandboxes: An Execution Layer for Agentic AI — wandb · 2026-08-01
- Engine-Agnostic Rust LLM Gateway SMG Graduates from LightSeek Foundation — vllm_project · 2026-08-01
- Benchmark: NInfer Boosts Qwen3.6 Prefill Speed Over 2x vs llama.cpp — tat_tvam_asshole · 2026-08-01