MoE expert expansion runs 35B Qwen on 16GB GPU at 60 tok/s, +2.5 pts on GPQA

Specific-Tax-6700 · reddit · 2026-10-05

A dev forked llama.cpp's server into AgrillaMoE, a dedicated engine for Qwen3.6-35B-A3B with Unsloth quants, hitting 57-60 tok/s on a rented 16GB V100 using the 2-bit UD-Q2KXL quant, exposing both OpenAI and Anthropic APIs so Claude Code works out of the box.

The key trick is runtime "MoE expansion": the model activates 8 routed experts per token, and a --moe-experts 20 patch with an adaptive threshold (keep experts with p ≥ 0.8 × rank-8 probability, applied to layers 25-39) consults more of the 35B parameters without retraining or file changes. Result: GPQA-Diamond rises from 81.82% (stock top-8) to 84.34% (+2.5 pts) on identical weights.

Original post →

More from Infra

Infra channel →