DeepSeek V4-Pro's stop-control issue persists across scaffolds; 100M token test shows prompts don't fix it
karminski3 · x · 2026-08-16
Blogger karminski3 conducted 18 independent long-horizon agentic coding tests on DeepSeek-V4-Pro-0813, with up to 50 tool calls each, combining 3 scaffolds and 2 stop prompts, repeated 3 times, totaling 100M tokens. Conclusion: changing scaffolds or prompts does not fix the stop-control issue, which is independent of overfitting. In tests, the model called finish only 3 times, continued after completion 4 times, and many runs lacked full verification.
Related event: DeepSeek V4-Pro Tests Reveal Reasoning and Control Issues(2 posts)→
More from coding & agent
- Collaborator: Infinite Canvas Environment for Agentic Development — tom_doerr · 2026-08-16
- Dev shares Hermes dual-agent workflow: Desktop and VPS collaboration — Teknium · 2026-08-16
- Stop Burning Tokens on Code Review: Use Custom Linters Instead — Elijah_Meeks · 2026-08-16
- Why Coding Agents Write Everything as a Captioning Task — keunwoochoi · 2026-08-16
- GitHub Copilot Deprecates Model, Sparking Debate on Vendor Lock-in — amu4biz · 2026-08-16
- Shipping code entirely remotely with agents, ditching local dev env — Register Spill (Thorsten Ball) · 2026-08-16