Qwen3.8-Flash-Next Cut 44% via REAP Hits 70% on Terminal-Bench 2.1
rmonsurate · reddit · 2026-10-01
The author used a lab Dell B300 to produce two open-weight fine-tunes of Qwen3.8-Flash-Next.
Victoria (coding and agents)
- Cut experts from 512 to 288 per layer with REAP, shrinking the model 44%
- Retrained at 4-bit (NVFP4) so it is trained for its shipping format rather than post-hoc quantized
- Terminal-Bench 2.1: 70.0% averaged over 3 runs (previous NVFP4 build: 62.5%); HumanEval 159/164
- 48.0 GiB of weights including the draft head; a separate 95.4 GiB n-gram table is not counted
- 280 tok/s single-stream on one B300 with the draft head vs 135 without
- GGUF Q4KM is 49.17 GiB, scoring 75.3% on Terminal-Bench (single run, noisy) and 93.2% on HumanEval
- Uses 35% fewer output tokens than the previous build
Maple (Canada-first)
Fine-tuned on top of Victoria so it defaults to Canada instead of the US for tax, benefits and regulatory questions. On 600 held-out questions with search: citing an official Canadian source rose from 6.0% to 62.9%; fully correct answers from 6.6% to 21.8%; "no answer" responses fell from 47.2% to 23.7%; pushing Canada onto users who said they live elsewhere dropped from 2.9% to 1.0%. HumanEval holds at 157/164.
The author notes llama.cpp users must build from the fork's qwen4exp-mtp branch, since mainline does not yet recognize the bundled draft head.
More from Models
- Can local LLMs handle Blender and game dev? Reddit says they still fall short — Any-Lingonberry7411 · 2026-10-01
- Rox benchmarks: Jev reranking beats GPT-5 Mini — 20x faster, 10x cheaper, 12% more accurate — hardimanjames · 2026-10-01
- Nat Lambert: More Frontier Labs Like Google's Gemini 4 Benefit Consumers — natolambert · 2026-10-01
- Google launches Gemini 4 Argon, a cybersecurity model that tops prompt injection benchmarks — ralucaadapopa · 2026-10-01
- Model wars: OpenAI went from best model in the world to arguably third place in a week — signulll · 2026-10-01
- Rumor: Gemini 4 spotted with SOTA knowledge scores, coding near Astra/Fable 5.1 level — haider1 · 2026-10-01