Early Looks: DeepSeek V4.1 Crushes Agentic/Coding Tasks but Regresses on Some Benchmarks
teortaxesTex · x · 2026-09-17
Early, unconfirmed impressions of DeepSeek V4.1 suggest it strongly outperforms the V4 series on agentic and coding tasks and is much larger than V4 Flash — yet it scores the same or worse on several high-signal benchmarks like MathArena and CritPt, making ARC-2 results the key test. Another developer who briefly benchmarked it found it roughly on par with DeepSeek V4 Flash before funding ran out.
More from Models
- Jev buzz reveals many don't know encoder-only classifiers have existed for years — RichmanRonald · 2026-09-17
- X's in-app Grok may soon require linking a standalone Grok account — nima_owji · 2026-09-17
- Quick test of OpenRouter's anonymous union-alpha model suggests it's no frontier flagship — karminski3 · 2026-09-17
- OpenRouter's mystery model union-alpha flagged as non-flagship by benchmark-less test — karminski3 · 2026-09-17
- Steve Sinofsky: Strip the anthropomorphism — model misbehavior is just bugs — ylecun · 2026-09-17
- Grok Bot DAU Jumps 80% in 5 Days, from ~67K to 121K, Similarweb Data Shows — FinanceYF5 · 2026-09-17