DeepSeek-V4-Flash reaches 105 tok/s on two 4090D cards after Triton kernel rewrites
iSevenDays · reddit · 2026-07-24
A Reddit user reports running DeepSeek-V4-Flash at about 105 tok/s on two Nvidia 4090D 48G cards with vLLM, after re-implementing Blackwell-only kernels in Triton.
Key details:
- Recreated DeepGEMM, FlashInfer sparse-MLA, and block-scaled FP8 paths for sm89 GPUs.
- Uses a patched setup on a Dell R740 with P2P enabled.
- Claims 2–3x better throughput for parallel agentic workloads and around 262k context.
- Includes commands for both the vLLM setup and a fully packed llama.cpp run.
The post is essentially a field report on squeezing modern model performance out of non-Blackwell hardware.
More from Infra
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27
- MiniBot 2.40 adds xAI, HF Studio and vLLM support with inline media tools — Creative-Type9411 · 2026-07-27
- Apple smart glasses, Nvidia-SK AI data center deal, and Ctrip’s RMB 5.179 billion fine headline a tech roundup — APPSO · 2026-07-27
- DeepSeek funding rumor, EU AI transparency rules and OpenAI agent incident make a packed AI news roundup — 创业邦 · 2026-07-27
- QuixiCore argues native quantized kernels beat dequant-then-generic execution — QuixiAI · 2026-07-27