Community miners make Qwen 27B FP8 3.5x faster than stock vLLM on a single H200
const_reborn · x · 2026-09-07
Pareton reports its mining community pushed Qwen3.8-27B-FP8 to 3.5x end-to-end speedup over stock vLLM on one H200 (median over SWE-agent traces):
- Median request latency 727ms → 190ms; throughput 58.7 → 221.8 tok/s; inter-token p99 93ms → 38ms; SLA-compliant requests 3% → 100%
- Outputs identical (≥0.99 greedy token-match)
- Winning patch: enabling the checkpoint's own MTP head for self-speculative decoding (γ=8) plus fused Triton kernels for GDN linear-attention layers — 1,090 lines inside vLLM
- Week 1: 25 rounds, 74 submissions, 31 miners
More from Infra
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11