antirez: DeepSeek v4.1 outscores Mistral Large 4 on DeepSWE 1.1 and other benchmarks
antirez · x · 2026-10-07
antirez points out that DeepSeek v4.1 scores higher than Mistral Large 4 on DeepSWE 1.1, and similar patterns hold on other benchmarks. He says he's happy to see more European LLMs but unhappy when comparisons cherry-pick the less obvious matchups.
More from Models
- Mistral Large 4 generates a Japanese-inspired floating voxel island, sparking 'Is the EU back?' buzz — kevinkern · 2026-10-07
- Marin 535B-A23B open model training crosses halfway, Percy Liang shares learnings — ericjang11 · 2026-10-07
- llama.cpp ships Day-0 support for Google's EmbeddingGemma 2 — ggerganov · 2026-10-07
- Trying to Plug Open-Source Mistral Into an Agentic Coder Just Doesn't Work, Says Berman — MatthewBerman · 2026-10-07
- StartLux claims its 27B model beats Jev AI on 31 of 38 benchmarks, self-reported results — Dr_Singularity · 2026-10-07
- 24 models tested on 669 clinical decisions: Jev stays #1 as two free models close in — MaziyarPanahi · 2026-10-07