Swift-Qwen3.8-27B GGUF with vision fits 262k context in 32GB VRAM at 76 tok/s
Then_Blueberry7290 · reddit · 2026-09-28
A local-deployment enthusiast highlights LuffyTheFox's Swift-Qwen3.8-27B-Genesis-GGUF: it includes vision yet comes in under 17GB, smaller than typical NVFP4 quants (19-20GB+), fitting full 262k context in 32GB VRAM with the vision projector offloaded to RAM.
On 2×5060 Ti 16GB OC with llama.cpp they benchmark 76 tok/s (range 40-113) at 262k context, and 45-65 tok/s in normal agentic workloads. Comparable NVFP4 models like thinkingcap only allow 160k context under vLLM and can't offload mmproj. Quality looks on par with other swift NVFP4 models; the poster asks what tradeoff the smaller size hides.
More from Infra
- Spectral deflation framework improves Muon: consistent validation loss gains in GPT-2 pretraining — hankyang94 · 2026-09-28
- First Audited Look at Inference-Economics: MiniMax Hit 24.6% Margin, Peer Lost 75% of OpenRouter Volume — AccBalanced · 2026-09-28
- Yunnan Germanium report: indium for InP is tight in China, export controls aren't the bottleneck — pstAsiatech · 2026-09-28
- Apple's free on-device fm paired with decision model Jev beats big-model routing in tests — jasonkneen · 2026-09-28
- Estimate: crudely describing human biology needs 1000x more data than humanity stores — IgorCarron · 2026-09-28
- Cooling setup runs 6x RTX 6000 at full 325W for 3 months, GPUs at just 41°C — TheZachMueller · 2026-09-28