1B+ Tokens Tested: Developer Accuses New DeepSeek Model of Benchmaxing
Sentdex · x · 2026-08-06
Developer @Sentdex shared his experience running DSV4F-Preview locally on over 1 billion tokens. While he loves the model, he feels the 0731 version is actually inferior to its predecessor.
He questions whether the reported benchmarks for 0731 are a case of extreme "benchmaxing" and asks the community: Does anyone actually feel 0731 is equivalent or better than GLM 5.2? Has anyone independently verified its scoring advantage in a given harness?
Related event: Developer Questions DeepSeek New Model Benchmark After Testing(2 posts)→
More from Models
- NVIDIA Launches Alpamayo 2 Super: Open Reasoning Model for Robotaxis — minchoi · 2026-08-06
- Hands-on: Muse Code Shines in 3D Visual Agentic Coding Task — cedric_chee · 2026-08-06
- DeepSeek V4 Flash Hits Local: Community Discusses VRAM Needs and Deployment — schaka · 2026-08-06
- LLMs for Quant Trading? Practitioners Express Skepticism — doodlestein · 2026-08-06
- Muse Spark 1.2 Hits 200 TPS in Tests, Priced at a Fraction of DeepSeek — alexandr_wang · 2026-08-06
- Maple 20B Hits 9,885 tokens/s on Single GH200 with 64 Concurrent Requests — MaziyarPanahi · 2026-08-06