DeepSeek New Model Tested: Great Long-Context, But Tool Calling Quirks

TheZachMueller · x · 2026-08-01

After overnight testing of the suspected DeepSeek v4-flash (0731) model, a developer shared positive initial impressions. However, they noted that the prompt format and tool call/conversation history quirks could cause integration issues, potentially leading to mixed early reviews—though some negative feedback might stem from user or harness errors.

In a subsequent coding agent gauntlet, the tester found the model's task persistence so robust that they had to raise its token limit from 200k to over 300k before it would give up on a task.

Related event: DeepSeek V4 Flash Tested: Great Long-Context, Tricky Integration(3 posts)→

Original post →

More from coding & agent

coding & agent channel →