Running 360k context DeepSeek on 4x RTX 3060 at ~100 tok/s

syscomua · reddit · 2026-08-18

A detailed technical report on running the DeepSeek V4 Flash Q4KXL model (144GB) across four RTX 3060 12GB GPUs.

By tweaking -ncmoe and tensor split parameters, the user achieved 99.4 tok/s prompt processing with a 368k context window. The post highlights that microbatch size (-ub) was the biggest performance lever and documents the specific VRAM constraints and layout strategies.

Original post →

More from Infra

Infra channel →