Dev runs Qwen3.8-Flash-Next on Quad 3090s with custom P2P driver, hitting 361.7 tok/s

QuixiAI · x · 2026-09-07

Developer QuixiAI reports running nvidia/Qwen3.8-Flash-Next-NVFP4 on Quad 3090s via SlimServe, a custom serving stack, achieving 120 tok/s at concurrency 1 and 361.7 tok/s at concurrency 8. The key enabler is their own open-source P2P kernel driver (QuixiAI/open-gpu-kernel-modules) for cross-GPU peer-to-peer communication, and they say optimization is ongoing.

Related event: Eric Hartford Forks NVIDIA Driver to Unlock Consumer GPU P2P, Open-Sources SlimServe(10 posts)→

Original post →

More from Infra

Infra channel →