Qwen3.6-35B hits 89.6% HumanEval on a single RTX 2080 Ti; MoE expert-expansion routing nudges it to 90.9%
Specific-Tax-6700 · reddit · 2026-10-07
The author ran Qwen3.6-35B-A3B with Unsloth's UD-IQ4XS dynamic 4-bit quant (fully resident in VRAM, no offload) on a single RTX 2080 Ti 22GB, scoring 89.6% pass@1 on the full original HumanEval (164 problems).
They then ran a controlled A/B of their "MoE expansion" inference-time routing patch — no retraining, no file changes. The router activates 20 experts per token instead of stock top-8 on layers 25–39 (with an adaptive threshold so easy tokens activate fewer), effectively consulting more of the network per token. Result: 90.9% (+2 problems) at −19% decode speed (69→56 tok/s). The same trick previously lifted GPQA-Diamond from 81.82% to 84.34% at Q8.
Honest caveats: +2/164 is within statistical noise, so read it as "equal or slightly better"; original HumanEval tests, not EvalPlus+, so no 1:1 leaderboard comparison; greedy n=1 protocol.
Also open-sourced AgrillaMoE, a llama.cpp-based server that auto-detects VRAM, downloads the right quant, applies the expansion profile by default, and exposes both OpenAI- and Anthropic-compatible APIs (Claude Code works out of the box) on NVIDIA GTX 10xx–RTX 50xx, AMD Vulkan, and Apple Silicon Metal.
More from Infra
- Ramjet: an open-source local alternative to NVIDIA Dynamo for multi-GPU inference — DoggoProfessor959 · 2026-10-07
- NVIDIA's NeMo-DCR Cuts Trillion-Parameter RL Weight Sync from 87.5 min to 150s — nvidia · 2026-10-07
- Qwen3.8 Flash Next GGUF benchmark: IQ3_S the sweet spot, 42.7M tokens tested — lxfater · 2026-10-07
- Used PS5 Pro hits $1,399 at GameStop as AI datacenters squeeze memory supply — aakashgupta · 2026-10-07
- Musk: xAI will build and run its Terafab itself, TSMC may only sublease part of it — MickeySteamboat · 2026-10-07
- "72% of the intelligence with 3.8% of the GPUs": Mistral's compute-efficiency ratio sparks debate — cyb3rops · 2026-10-07