DeepSeek V4 Flash Hits 1M Context Locally

Shoddy_Bed3240 · reddit · 2026-07-18

A Reddit user shared their configuration and test results running DeepSeek V4 Flash on a 5090 using llama.cpp, highlighting that they pushed the context length to 1 million tokens.

They used the GGUF version provided by Unsloth and posted the full llama-server startup parameters, including:

Measured performance is roughly:

The author believes the speed isn't as ideal as the Qwen series yet, but notes there is still room for optimization in llama.cpp.

Original post →

More from Infra

Infra channel →