DeepSeek V4.1 Flash scores 24/105 for $1.80 in community eval test
PawelHuryn · x · 2026-09-10
Blogger PawelHuryn benchmarked new DeepSeek models on his eval suite: DeepSeek-V4.1-Flash at max effort scored 24/105 for $1.80 in 42.6 min (19/105 at $0.31 on high effort). Comparisons: Gemini 3.8 Flash (high) 20/105 at $9.78; GLM-5.3 19/105 at $19.73 in 66.7 min. A follow-up test of DeepSeek V4 Pro (max) scored 16/105 at $1.89 vs V4.1 Flash (max) 24/105 at $1.08—significant gains at far lower cost. Updates tracked on GitHub.
Related event: Bug Hunt Bench: DeepSeek V4.1 Flash Tops Price-Performance(8 posts)→
More from Models
- Anthropic Accuses Moonshot of Routing 300K User Queries to Claude via 5,380 Fake Accounts — toptickcrypto · 2026-09-11
- DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture — ccerrato147 · 2026-09-11
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Assistant Benchmark goes live: 61 assistants scored across 15 real-use dimensions — Scobleizer · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11