DeepSeek V4 local tool calling fails, suspected feedback loop

No-Paper-557 · reddit · 2026-08-21

A user reports significant issues deploying DeepSeek-V4-Flash as a local coding agent on a single RTX PRO 6000 Blackwell (96GB). Despite successfully loading the model using a custom vLLM-MoE build, continuous tool calling tests revealed critical failures:

The user suspects the cause lies in the V4 encoder's mechanism, which retains and renders previous assistant reasoning in later turns. This may create a feedback loop where the model confuses past plans with current results. Similar upstream vLLM reports regarding DSML leakage were also noted.

Original post →

More from coding & agent

coding & agent channel →