Qwen 3.8 27B NVFP4 benchmarked on 2xV100 with ~400k context
jjusko20 · reddit · 2026-09-17
A Reddit user benchmarks the NVFP4-quantized Qwen 3.8 27B on 2x V100 32GB with NVLink, using dflash2 to reach roughly 400k total context, with a screenshot of results. Useful reference for running long-context local inference on older GPUs.
More from Infra
- How to run Qwen3.8-Flash-Next with N-gram SSD streaming in llama.cpp? — Ambitious_Fold_2874 · 2026-09-17
- How Bell Labs Missed the Microchip: IEEE Spectrum Revisits a Landmark Tech-History Blunder — ArtificialOther · 2026-09-17
- Agentic AI systems are the next network users: 40% of enterprise apps to include agents by 2026 — seankinneyRCR · 2026-09-17
- B200 spot rental up 80% in 8 months as demand outpaces compute buildout — JOBhakdi · 2026-09-17
- MLX-Serve 26.9.3 ships: Qwen Flash Next tops 100 tok/s on M4/M5 Max Macs — TheMoonMidas · 2026-09-17
- Running a 14B Model on 16GB RAM: 'My PC Is a Toaster Now' — Aggravating_Site381 · 2026-09-17