Swift 1.5 + HyperQwen cuts task time 37% at 100+ tok/s on a single RTX 3090
KingGongzilla · reddit · 2026-09-29
A community benchmark adapts UkisAI's Swift 1.5 (Qwen3.8 27B finetune) to the HyperQwen serving stack: on one RTX 3090 24GB with FP8 KV cache and 150k context, average time per task drops 37% (108.1s → 68.2s) across 630 benchmark tasks, while decode speed stays above 100 tok/s thanks to fewer generated tokens. Quality holds up: GSM8K 97.5%, LiveCodeBench 91%, 30/30 on a custom tool-call/JSON eval.
- INT8-head variants quantize the output head and MTP layers, adding HyperQwen's draft vocabulary for speculative decoding
- The 'fast' variant uses GPTQ INT4 heads plus a purpose-built 65,536-token draft vocabulary
- No additional finetuning was needed; weights and RUNTIME.md setup instructions are public on Hugging Face
More from Infra
- Grass founder: training data remains the first product for frontier AI labs — toptickcrypto · 2026-09-29
- Modal Labs closing in on $750M round at $15.75B valuation — TechCrunch AI · 2026-09-29
- exe.dev kills per-seat pricing, shifts to compute pools; individual plan drops to $15/mo — davidcrawshaw · 2026-09-29
- Exo cuts base plan to $15/month for 2 vCPUs / 4 GB, swaps included tokens for BYO ChatGPT subscription — davidcrawshaw · 2026-09-29
- Open-weight 0.8B/2B System 1 decision models match Jev at 83.1%, trained fully on local hardware — Usual_Maximum7673 · 2026-09-29
- Qwen3.8 Flash hits 74 tok/s single-stream, 212 tok/s aggregate on one DGX Spark — open vLLM recipe — DimeRhyme · 2026-09-29