1B+ Tokens Tested: Developer Accuses New DeepSeek Model of Benchmaxing

Sentdex · x · 2026-08-06

Developer @Sentdex shared his experience running DSV4F-Preview locally on over 1 billion tokens. While he loves the model, he feels the 0731 version is actually inferior to its predecessor.

He questions whether the reported benchmarks for 0731 are a case of extreme "benchmaxing" and asks the community: Does anyone actually feel 0731 is equivalent or better than GLM 5.2? Has anyone independently verified its scoring advantage in a given harness?

Related event: Developer Questions DeepSeek New Model Benchmark After Testing(2 posts)→

Original post →

More from Models

Models channel →