Swift 1.5 + HyperQwen cuts task time 37% at 100+ tok/s on a single RTX 3090

KingGongzilla · reddit · 2026-09-29

A community benchmark adapts UkisAI's Swift 1.5 (Qwen3.8 27B finetune) to the HyperQwen serving stack: on one RTX 3090 24GB with FP8 KV cache and 150k context, average time per task drops 37% (108.1s → 68.2s) across 630 benchmark tasks, while decode speed stays above 100 tok/s thanks to fewer generated tokens. Quality holds up: GSM8K 97.5%, LiveCodeBench 91%, 30/30 on a custom tool-call/JSON eval.

Original post →

More from Infra

Infra channel →