27B at Q5 with full 131k context on one 24GB RTX 3090, 13-17% faster
bjivanovich · reddit · 2026-10-01
A Reddit user released ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP, a GGUF quantization suite (i1-Q5KM primary) that runs a 27B model with the full 131,072-token context on a single RTX 3090 at 50-65+ t/s.
The challenge
- Standard Q5KM weights are 19.2GB; an FP16 KV cache at 131k tokens alone exceeds 24GB
- Community quants compress the MTP draft block uniformly, dropping speculative acceptance from 85% to 55%
The approach
- Asymmetric per-layer quantization: MTP head (blk.64) isolated at Q80 (acceptance back to 76-88%), attention layers protected at Q6K/Q80 to prevent long-context reasoning decay, FFN at Q5KM
- Unified Turbo KV cache (-ctk turbo5 -ctv turbo4) shrinks the 131k KV cache to 3.8GB, fitting everything fully in 24GB VRAM (-ngl 99)
Benchmarks (80k-98k context vs standard community Q5)
- Sustained speed 50-54 t/s vs 45 t/s (+13-17%)
- Peak bursts 10.7 t/s higher
- MTP draft acceptance +11.5-16.6%
Available on Hugging Face in i1-Q80/Q6K/Q5KM/Q4KM plus mmproj-BF16 for vision.
More from Infra
- Comfy API goes live: deploy workflow JSONs as autoscaling GPU endpoints — PurzBeats · 2026-10-01
- E2B open-sources Embed, packaging full agent sandbox stack on a single node — badphilosopher · 2026-10-01
- ByteDance Seed finds phase sensitivity in chunked KV-cache compression, retrieval accuracy swings 40 points — ByteDance-Seed · 2026-10-01
- HPE/Broadcom Tomahawk trays for AMD Helios racks: six per rack, copper-heavy scale-up — BenBajarin · 2026-10-01
- The most advanced part of ASML machines, the light source, is made in San Diego — Darpinian · 2026-10-01
- Agent plugin /who-ate-my-flops speeds PyTorch jobs up to 3.6x, PRs merged into six OSS repos — ruilong_li · 2026-10-01