Dev runs Qwen3.8-Flash-Next on Quad 3090s with custom P2P driver, hitting 361.7 tok/s
QuixiAI · x · 2026-09-07
Developer QuixiAI reports running nvidia/Qwen3.8-Flash-Next-NVFP4 on Quad 3090s via SlimServe, a custom serving stack, achieving 120 tok/s at concurrency 1 and 361.7 tok/s at concurrency 8. The key enabler is their own open-source P2P kernel driver (QuixiAI/open-gpu-kernel-modules) for cross-GPU peer-to-peer communication, and they say optimization is ongoing.
More from Infra
- Yacine wonders when AI-designed chips will be taped out directly for businesses and consumers — yacineMTB · 2026-09-07
- Yacine predicts PCB design will run on open-weight models on consumer multi-GPU rigs within a year — yacineMTB · 2026-09-07
- Leak: NVIDIA Vera SOCAMM De-Spec Remains, 64GB Only by Q1 2027, 4-Hi GPU Variants Unlikely — zephyr_z9 · 2026-09-07
- LayerStoRm open-source engine runs 186GiB MoE on 96GB VRAM at 24.5 tok/s with 1M context — CharacterBumblebee99 · 2026-09-07
- AI buildout has created 300k+ construction jobs since 2022, electrician and HVAC trades booming — soumitrashukla9 · 2026-09-07
- VRAMWATCH tracks live GPU street prices from RTX 5090 to H200 — KyeGomezB · 2026-09-07