Third-Party Benchmark Claims DeepSeek V4.1 Flash Matches 98% of GPT-6 at 1.4% of Cost
ChrisUniverse · x · 2026-09-10
OpenDesignHQ benchmarked DeepSeek V4.1 Flash on everyday design tasks, claiming it hit 98% of GPT-6 Astra's score at 1.4% of the cost, with every other model scoring lower while costing more. A repost adds $0.023/task, 3x faster, 5% higher score. Unverified: neither model name has official confirmation, so treat with skepticism.
More from Models
- DeepSeek Ships V4.1-Flash: 552B MoE With 8B Active, 75% Less KV Cache Memory — mark_k · 2026-09-11
- Grok Voice Think Fast 2.0 High tops speech-to-speech leaderboard on task success — XFreeze · 2026-09-11
- Fable and Astra fail at the XY problem: eager executors with zero pushback and no metacognition — teodorio · 2026-09-11
- 'Chart Crime': Blogger Flags Cropped TerminalBench 4.0 Chart Underrating DeepSeek V4.1 — teortaxesTex · 2026-09-11
- Real API pricing method ranks subscriptions: GPT-5.6 Luna cheapest at $0.00083/MTok — teortaxesTex · 2026-09-11
- ThursdAI weekly: Astra, Navier Stokes, Muse assistant and more AI news — thursdai_pod · 2026-09-10