MoE expert expansion runs 35B Qwen on 16GB GPU at 60 tok/s, +2.5 pts on GPQA
Specific-Tax-6700 · reddit · 2026-10-05
A dev forked llama.cpp's server into AgrillaMoE, a dedicated engine for Qwen3.6-35B-A3B with Unsloth quants, hitting 57-60 tok/s on a rented 16GB V100 using the 2-bit UD-Q2KXL quant, exposing both OpenAI and Anthropic APIs so Claude Code works out of the box.
The key trick is runtime "MoE expansion": the model activates 8 routed experts per token, and a --moe-experts 20 patch with an adaptive threshold (keep experts with p ≥ 0.8 × rank-8 probability, applied to layers 25-39) consults more of the 35B parameters without retraining or file changes. Result: GPQA-Diamond rises from 81.82% (stock top-8) to 84.34% (+2.5 pts) on identical weights.
More from Infra
- One USB-C cable turns an iPhone into a 24GB MacBook's extra memory, running local Qwen 27B 40%+ faster — TheMoonMidas · 2026-10-05
- Hugging Face Accelerate lead touts 6 years of inference work, v0.20.0 features — TheZachMueller · 2026-10-05
- Compute deals insider: buyers want only NVIDIA gear, CUDA moat alive and well — sudoraohacker · 2026-10-05
- Micron CEO: memory supply will be much tighter in 2027 and 2028 than in 2026 — chillinewman · 2026-10-05
- LFM2.5 2.6B beats MiniCPM5 2B in speed and RAM: 22 t/s vs 16 t/s on M1 Air — parepeg · 2026-10-05
- Indie dev's cheapest stack: Firebase analytics + Cloudflare R2 + Vercel — jdluk87 · 2026-10-05