DeepSeek V4 Flash Hits 1M Context Locally
Shoddy_Bed3240 · reddit · 2026-07-18
A Reddit user shared their configuration and test results running DeepSeek V4 Flash on a 5090 using llama.cpp, highlighting that they pushed the context length to 1 million tokens.
They used the GGUF version provided by Unsloth and posted the full llama-server startup parameters, including:
- Context length of 1,048,576
- K/V cache quantization
- Various CPU/GPU allocation and batch parameters
- Specifying chat template, repetition penalty, and service port
Measured performance is roughly:
- Prefill around 650–700 tok/s
- Decode around 17 tok/s
- Load time of about 32 seconds
The author believes the speed isn't as ideal as the Qwen series yet, but notes there is still room for optimization in llama.cpp.
More from Infra
- Nebius says SlimSpec speeds speculative decoding 8–9% without shrinking the vocabulary — Arindam_1729 · 2026-07-21
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- Chamath says open-sourcing Grok would push AI margins from models to infra and apps — Dan_Jeffries1 · 2026-07-21
- EU AI competitiveness is under pressure as firms double down on chips, ethics, and talent — nordicinst · 2026-07-21
- AI bottlenecks are shifting to memory, optics, yield control and power — thedealdirector · 2026-07-21