Dev Slams DeepSeek V4 Flash: Benchmark Scores Don't Match Real Coding Performance

adellknudsen · reddit · 2026-08-01

A developer expressed strong disappointment with DeepSeek V4 Flash on Reddit. Despite the model's sky-high benchmark scores, it performs poorly in real-world C/C++ programming tasks and even fails at simple Playwright automation.

The author argues that the trendy 'single-file HTML 3D demos' have become a visual standard for hype, but the underlying public code is likely already in the training data. He calls for a new, dynamic benchmark system with unknown tests to prevent open-source models from over-fitting to leaderboards. He also pushes back on the 'cheap model' cost-saving argument, noting that compensating for lower capability wastes too much developer time in 'vibe coding'.

Original post →

More from coding & agent

coding & agent channel →