Swift-Qwen3.8-27B GGUF with vision fits 262k context in 32GB VRAM at 76 tok/s

Then_Blueberry7290 · reddit · 2026-09-28

A local-deployment enthusiast highlights LuffyTheFox's Swift-Qwen3.8-27B-Genesis-GGUF: it includes vision yet comes in under 17GB, smaller than typical NVFP4 quants (19-20GB+), fitting full 262k context in 32GB VRAM with the vision projector offloaded to RAM.

On 2×5060 Ti 16GB OC with llama.cpp they benchmark 76 tok/s (range 40-113) at 262k context, and 45-65 tok/s in normal agentic workloads. Comparable NVFP4 models like thinkingcap only allow 160k context under vLLM and can't offload mmproj. Quality looks on par with other swift NVFP4 models; the poster asks what tradeoff the smaller size hides.

Original post →

More from Infra

Infra channel →