DeepSeek-V4-Flash Halts Mid-Task During Long Agentic Coding at 100K+ Tokens

dieSpaghettiCarbona · reddit · 2026-08-10

A developer reported an anomaly when running DeepSeek-V4-Flash-0731 (Q8KXL GGUF) locally via Unsloth Studio for long-running agentic coding tasks.

Once the context exceeds 100K tokens, the model occasionally stops generating mid-task without any explicit error. Typing resume prompts the model to correctly pick up where it left off, but the stopping issue recurs as the context grows. The developer speculates this could be tied to llama.cpp inference, context handling, prompt caching, or OpenCode, and asked the community if others are experiencing this at large context lengths.

Original post →

More from coding & agent

coding & agent channel →