Running 360k context DeepSeek on 4x RTX 3060 at ~100 tok/s
syscomua · reddit · 2026-08-18
A detailed technical report on running the DeepSeek V4 Flash Q4KXL model (144GB) across four RTX 3060 12GB GPUs.
By tweaking -ncmoe and tensor split parameters, the user achieved 99.4 tok/s prompt processing with a 368k context window. The post highlights that microbatch size (-ub) was the biggest performance lever and documents the specific VRAM constraints and layout strategies.
More from Infra
- DSCO Router Launches Unified Gateway for Multi-Model Routing with BYOK Support — arthurcolle · 2026-08-24
- Open Source RobotSoul: Persistent Identity for Agents After Context Resets — robauto-dot-ai · 2026-08-24
- Offloading MoE models to RAM causes slow prefill speeds — former_farmer · 2026-08-24
- Etched Raises $1B Led by Jane Street to Validate Architecture-Agnostic AI Chips — TheTuringPost · 2026-08-24
- ConvRot Quant joins llama-cpp: Q6 accuracy nears Q8 quality — giveen · 2026-08-24
- LifeOS: A Local, Voice-Driven Personal Organizer — Extension-Bid-639 · 2026-08-24