MTP vs MTP+Ngram on Qwen3.8 Flash: 10% Speed but 3x Token Usage
esw123 · reddit · 2026-09-03
Testing Qwen3.8 Flash Next in Unsloth Studio, the poster finds MTP alone gives 10% higher throughput but token usage jumps from 20K to 66K due to extra thinking. They ask whether Ngram should always be enabled and how to trade off speed, cost, and quality.
More from Infra
- Broadcom guides AI revenue to ~$115B in FY2027, doubling again to $230B in FY2028 — BenBajarin · 2026-09-03
- Perplexity open-sources Lily, its Apple Silicon inference engine for Qwen3.6-35B-A3B — inductionheads · 2026-09-03
- Nvidia now ~8% of S&P 500 market cap, worth 16.3% of US GDP at $5.3T — ivan_bezdomny · 2026-09-03
- Broadcom Q3 AI chip revenue hits $16.7B, up 221% YoY; guides $21.7B for Q4 — Beth_Kindig · 2026-09-03
- PINNACLE: Claude Fable 5.1 halves agent failure rate but costs $2.46 per correct answer — ryanshrout · 2026-09-03
- eBPF looks cheap but isn't free: Bitbison's deep dive on hooks, CO-RE reads and rings — tianyin_xu · 2026-09-03