Optimizing DeepSeek on RTX 3090: 128K Context Inference Benchmarks
Ok_Ninja7526 · reddit · 2026-08-06
A developer shared an in-depth optimization experiment on Reddit for running the DeepSeek-V4-Flash-0731 GGUF model on a single RTX 3090 (24GB), with a strict requirement to maintain a 128,000-token context window.
Using the llama.cpp backend, the author tested four quantization variants by tweaking GPU offloading, CPU expert placement, and KV-cache quantization:
- UD-IQ4XS: 9.9 tok/s, pushing system RAM and VRAM to their absolute limits.
- UD-IQ3S: 12.1 tok/s, achieving a 22% speed boost and reducing RAM usage by 17GB compared to IQ4XS.
- UD-IQ3XXS: 12.5 tok/s, showing only marginal improvement over IQ3S.
The post details the exact hardware specs (AMD Ryzen 9 9900X, 128GB DDR5) and software configurations, providing valuable engineering data for local deployment of long-context models.
More from Infra
- Running 35B Model on RTX 3090: Tuning MoE Offload Spikes Processing Speed by 2.36x — Longjumping-Music638 · 2026-08-06
- Jensen Huang Calls Algorithms the New Asset Class; Space Data Center Startup Starcloud Hits $1B Valuation — santoshpanda · 2026-08-06
- Hugging Face Hits Record 4PB Weekly Uploads as AI Treats Humans as Storage — vanstriendaniel · 2026-08-06
- Run Multiple Models on One GPU: SIE Cuts Self-Hosting Costs 75% — Roger_M_Taylor · 2026-08-06
- US Hyperscalers Set to Invest Over $700 Billion in AI Computing by 2026 — coinfanking · 2026-08-06
- Nebius Inference Platform Hits Milestone in Artificial Analysis Accuracy Index — demian_ai · 2026-08-06