Qwen3.8-Flash-Next on SGLang NVFP4: 254K-Context TTFT Drops 35s→22s, Full Gauntlet Tested
FantasticNature7590 · reddit · 2026-09-17
- On a single RTX PRO 6000 Blackwell (96GB), the author benchmarked Qwen3.8-Flash-Next with SGLang's official NVFP4 image: TTFT on a 254K-token prompt fell from 34.9s to 22.4s, prefill rose from 7,284 to 11,352 tok/s (1.56×), and cached repeat prompts dropped TTFT from 21.7s to 0.44s (49×).
- Capability results: long-context recall 81/81 (fully correct even at 99% window fill); BFCL single-turn tool accuracy 85.2% vs 49.5% multi-turn; τ²-bench telecom 68.1% completion but only 41.2% passed all three attempts; battle gauntlet 13/17 wins with thinking on, 8/17 off.
- SVG, video-editing and animated-design generations mostly passed hard checks (video editing scored 10/10 on the author's rubric). Full hardware config, methodology and latency/throughput tables across context lengths included — a rare system-level long-context benchmark on an open stack.
More from Infra
- Common Crawl puts crawl archives on Hugging Face Storage Bucket, with a getting-started guide — vanstriendaniel · 2026-09-17
- Google brings Gemma 4 12B to Mac, running locally on 16GB MacBook Air — GlennCameronjr · 2026-09-17
- Weekly token usage up 25,000%+ in ~600 days, room for many winners — brucemacv · 2026-09-17
- Zilliz CTO outlines 'One Data, One Index' architecture to reshape agent retrieval — J_Luan_ · 2026-09-17
- Xiaomi livestreams MIMO-V2.6 post-training, burning $10 per second — Dr_Karminski · 2026-09-17
- Retirees should buy an Nvidia Spark and local models to capture a lifetime of ideas — TinfoilTricorn · 2026-09-17