Qwen3.8-Flash-Next-NVFP4 hits 120-361 tok/s on Quad 3090 via custom P2P driver
QuixiAI · x · 2026-09-07
QuixiAI got nvidia/Qwen3.8-Flash-Next-NVFP4 running in SlimServe on a quad RTX 3090 setup, reaching about 120 tok/s at concurrency 1 and 361.7 tok/s at concurrency 8, using a custom P2P driver (QuixiAI/open-gpu-kernel-modules).
More from Infra
- 1-bit quantized embeddings cut vector index storage up to 60x with <1% quality loss — burkov · 2026-09-07
- Starlink caps heavy 'unlimited' users to 10Mbps after ~5TB monthly usage — mcraddock · 2026-09-07
- lauriewired: CXL.mem via a good switch at ~500ns should still beat 1-4us RDMA — lauriewired · 2026-09-07
- Anthropic commits $517B in compute deals, ~3x what it told investors — GaryMarcus · 2026-09-07
- SemiAnalysis drops 52-minute deep dive: AI is running out of power — dylan522p · 2026-09-07
- Not enough CPUs to run all agents: sandbox architecture needs a rethink, says Anthropic exec — dinasaur_404 · 2026-09-07