DeepSeek V4-Pro's stop-control issue persists across scaffolds; 100M token test shows prompts don't fix it

karminski3 · x · 2026-08-16

Blogger karminski3 conducted 18 independent long-horizon agentic coding tests on DeepSeek-V4-Pro-0813, with up to 50 tool calls each, combining 3 scaffolds and 2 stop prompts, repeated 3 times, totaling 100M tokens. Conclusion: changing scaffolds or prompts does not fix the stop-control issue, which is independent of overfitting. In tests, the model called finish only 3 times, continued after completion 4 times, and many runs lacked full verification.

Related event: DeepSeek V4-Pro Tests Reveal Reasoning and Control Issues(2 posts)→

Original post →

More from coding & agent

coding & agent channel →