27B at Q5 with full 131k context on one 24GB RTX 3090, 13-17% faster

bjivanovich · reddit · 2026-10-01

A Reddit user released ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP, a GGUF quantization suite (i1-Q5KM primary) that runs a 27B model with the full 131,072-token context on a single RTX 3090 at 50-65+ t/s.

The challenge

The approach

Benchmarks (80k-98k context vs standard community Q5)

Available on Hugging Face in i1-Q80/Q6K/Q5KM/Q4KM plus mmproj-BF16 for vision.

Original post →

More from Infra

Infra channel →