DeepSeek V4 Review: Frontier at CVE Finding, Weak on Hallucination
teortaxesTex · x · 2026-08-26
Provides a detailed multi-dimensional assessment of the DeepSeek V4 model:
- Strengths: Frontier performance at CVE finding/recal; open frontier on ProgramBench-Vetted, ARC-2, and potentially math (LiveBench); catching up on CritPt.
- Weaknesses: Mediocre on all sorts of SWE and agentic evals; garbage on hallucination, reward hacking, and design taste.
The author concludes with a suggestive question about the model's nature.
More from Models
- Test Shows Flash-Vision-Excels at Kernel Dev but Fails Logic Integration — teortaxesTex · 2026-08-26
- Only Grok Correctly Answers 'Rich Get Richer' Logic Trap, Others Fail — davidpattersonx · 2026-08-26
- Ox Alpha confirmed as Zhipu GLM-5.3-Flash with 1M context — MrWidmoreHK · 2026-08-26
- Claude Opus 5 works okay only if you set autocompact to 200k tokens, dev reports — bclavie · 2026-08-26
- User complaints about new GPT models' hallucinations and slowness — SweatyActuator2119 · 2026-08-26
- User reports ChatGPT Pro 'dumbing down' issue fixed — mazzaTalk · 2026-08-26