Early Looks: DeepSeek V4.1 Crushes Agentic/Coding Tasks but Regresses on Some Benchmarks

teortaxesTex · x · 2026-09-17

Early, unconfirmed impressions of DeepSeek V4.1 suggest it strongly outperforms the V4 series on agentic and coding tasks and is much larger than V4 Flash — yet it scores the same or worse on several high-signal benchmarks like MathArena and CritPt, making ARC-2 results the key test. Another developer who briefly benchmarked it found it roughly on par with DeepSeek V4 Flash before funding ran out.

Original post →

More from Models

Models channel →