DeepSeek V4 tested: ARC-AGI reasoning costs drop despite higher performance
DeArgonaut · reddit · 2026-08-09
A developer tested DeepSeek V4 (0731) on the ARC-AGI-1 and ARC-AGI-2 benchmarks. The results show that this version breaks the industry trend where performance gains come with increased reasoning costs, achieving higher scores while actually decreasing the cost per task.
More from Models
- Leak: OpenAI's Largest Pre-Train 'Doug' Coming Year-End — ChrisGPT · 2026-08-09
- Researcher Asks About HF Poisoning: Were Corrupted Checkpoints Reverted? — jd_pressman · 2026-08-09
- ChatGPT Swears But Claude Stays Professional: User Notes Tone Differences — ignorantwat99 · 2026-08-09
- Liquid AI Releases LFM 2.5: A 2.6B Parameter Model — JosephJacks_ · 2026-08-09
- Liquid AI's New Model Outperforms 4x Larger Peers in Private Benchmarks — JosephJacks_ · 2026-08-09
- Experiment Shows Naive LLM Code-Auditing Loops Destroy Correct Code — sebpaquet · 2026-08-09