DeepSeek V4 Hits 1,328 tok/s Prefill on Single RTX 6000 via Krasis
mrstoatey · reddit · 2026-08-04
Using the inference optimization tool Krasis (v1.0.19), a developer achieved significant speedups for DeepSeek-V4-Flash-0731 (INT4) on a single RTX PRO 6000 (96GB).
By streaming the model across VRAM and system RAM, the tool keeps 6,440 of the hottest experts resident while offloading the rest. In tests, prefill speeds reached an impressive 1,328 tok/s for a 23k token prompt, with decode speeds of 29.4 tok/s at a 1K context. This offers a highly cost-effective local deployment setup for coding agents that frequently send massive contexts.
More from coding & agent
- Snorkel AI Builds Simulated Enterprise Environments to Train AI Agents on Complex Workflows — ajratner · 2026-08-05
- Minnow: Open-Source AI Workspace Integrating Chat, Deep Research, and Multi-Agent Orchestration — MinnowAI · 2026-08-05
- Goodfire Launches Silico Platform for Frontier-Scale Model Interpretability and Training — Jeande_d · 2026-08-05
- Grok Integrated into GitHub Copilot: Testing its App Building Capabilities — DanWahlin · 2026-08-05
- AI Agent Installs Dual-Boot System in Seconds for Just $0.01 — yacineMTB · 2026-08-05
- Podcast Explores Personal AI Agent Development: From Solving Self Needs to Shipping — msg · 2026-08-05