Qwen3.6-35B hits 89.6% HumanEval on a single RTX 2080 Ti; MoE expert-expansion routing nudges it to 90.9%

Specific-Tax-6700 · reddit · 2026-10-07

The author ran Qwen3.6-35B-A3B with Unsloth's UD-IQ4XS dynamic 4-bit quant (fully resident in VRAM, no offload) on a single RTX 2080 Ti 22GB, scoring 89.6% pass@1 on the full original HumanEval (164 problems).

They then ran a controlled A/B of their "MoE expansion" inference-time routing patch — no retraining, no file changes. The router activates 20 experts per token instead of stock top-8 on layers 25–39 (with an adaptive threshold so easy tokens activate fewer), effectively consulting more of the network per token. Result: 90.9% (+2 problems) at −19% decode speed (69→56 tok/s). The same trick previously lifted GPQA-Diamond from 81.82% to 84.34% at Q8.

Honest caveats: +2/164 is within statistical noise, so read it as "equal or slightly better"; original HumanEval tests, not EvalPlus+, so no 1:1 leaderboard comparison; greedy n=1 protocol.

Also open-sourced AgrillaMoE, a llama.cpp-based server that auto-detects VRAM, downloads the right quant, applies the expansion profile by default, and exposes both OpenAI- and Anthropic-compatible APIs (Claude Code works out of the box) on NVIDIA GTX 10xx–RTX 50xx, AMD Vulkan, and Apple Silicon Metal.

Original post →

More from Infra

Infra channel →