DeepSeek-V4.1-Flash aces 22 knowledge-work tasks for just $0.34 in eval run
realsohamparekh · x · 2026-09-10
A practitioner evaluated the newly released DeepSeek-V4.1-Flash on 22 internal knowledge-work tasks (engineering, finance, healthcare), using 120+ verifiers on average with GPT-5.5 as judge — the entire run cost $0.34.
Strengths: applying long, interacting policies with precedence and effective dates; reconciling multiple CSVs, docs and amendments into consistent outputs; quantitative/financial reasoning; large contexts; passing hidden semantic generalization tests.
Weaknesses: distinguishing primary vs. secondary findings in summary counts; occasionally solving the underlying audit but omitting one required comparison in exhaustive memos.
The post quotes DeepSeek's official launch: V4.1-Flash is the smallest model in its new architecture family, with native visual understanding and a focus on faster inference and higher throughput. Third-party numbers are unverified by DeepSeek.
More from Models
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11
- Claims resurface that Moonshot's Kimi distilled from Claude raw CoTs — xuanalogue · 2026-09-11
- User switches back to GPT-5.6 Sol: barely uses quota and feels faster — CtrlAltDwayne · 2026-09-11
- Dev opinion: model differences shrink in a good harness; Grok 4.6 is good enough — gnukeith · 2026-09-11