DeepSeek evals v4.1 flash in 8 harness configs — best in minimal harnesses, worse in Claude Code and Codex
zainhas · x · 2026-09-10
DeepSeek benchmarked its new v4.1 flash model across 8 different agent harness configurations. The model performs best in minimal harnesses (mini-SWE and the minimal deepseek harness) but underperforms in both Claude Code and Codex — a sign its fit with heavier harnesses still needs work.
Related event: DeepSeek Releases Harness v0.1.5, Evaluates V4.1 Flash Across Configs(2 posts)→
More from coding & agent
- Open-source TIDAL MCP server ships 112 tools for music discovery via Claude and Cursor — Fickle_Guitar7417 · 2026-09-10
- CommerceAgentBench hits 1,000 GitHub stars with 107 tasks to vet e-commerce agents — VibeMarketer_ · 2026-09-10
- 8 AI tools gaining tons of GitHub stars this week: Skills, Archify, MiniMind and more — Shruti_0810 · 2026-09-10
- mcp-x: 42 X API tools designed so the model can't burn your money — GoldBroccoli7073 · 2026-09-10
- Replace RPA with Doubao's browser record-and-replay: automate link submissions — lxfater · 2026-09-10
- DeepSeek's new open model beats GLM 5.3 and Kimi K3 at 4-10x lower price — deedydas · 2026-09-10