DeepSeek-V4.1-Flash aces 22 knowledge-work tasks for just $0.34 in eval run

realsohamparekh · x · 2026-09-10

A practitioner evaluated the newly released DeepSeek-V4.1-Flash on 22 internal knowledge-work tasks (engineering, finance, healthcare), using 120+ verifiers on average with GPT-5.5 as judge — the entire run cost $0.34.

Strengths: applying long, interacting policies with precedence and effective dates; reconciling multiple CSVs, docs and amendments into consistent outputs; quantitative/financial reasoning; large contexts; passing hidden semantic generalization tests.

Weaknesses: distinguishing primary vs. secondary findings in summary counts; occasionally solving the underlying audit but omitting one required comparison in exhaustive memos.

The post quotes DeepSeek's official launch: V4.1-Flash is the smallest model in its new architecture family, with native visual understanding and a focus on faster inference and higher throughput. Third-party numbers are unverified by DeepSeek.

Related event: DeepSeek V4.1 Flash tested across tasks: near-frontier performance at a fraction of the cost(15 posts)→

Original post →

More from Models

Models channel →