DeepSeek V4.1 model card shows same model scores wildly differently across agent harnesses
solyarisoftware · x · 2026-09-12
One of the most interesting tables in the DeepSeek V4.1 model card shows the same model, on the same benchmark, scoring wildly differently depending on the agent harness used. The poster argues this proves that comparing a new model's scores against previously published numbers is close to meaningless unless the harness and config are matched — a fresh reminder that leaderboard numbers are heavily scaffold-dependent.
More from Models
- Codex quality fixes shipped, reset rolling out to ChatGPT Work users — CtrlAltDwayne · 2026-09-12
- Codex usage reset reportedly rolling out Sep 12 at 07:00 UTC, execution unconfirmed — DevDminGod · 2026-09-12
- Qwen3.8-27B goes live on Cerebras with fast inference, scoring 34 on AAII — Alibaba_Qwen · 2026-09-12
- antirez uploads DeepSeek V4.1 Flash GGUF quantizations to Hugging Face — Queasy_Asparagus69 · 2026-09-12
- Qwen3.8-Max scores 19 on OpenRouter vs 26 on Alibaba — effort level bug suspected — PawelHuryn · 2026-09-12
- Pipecat v1.9 adds Meta's Muse Voice Transcribe, the lowest semantic-WER STT model tested — solyarisoftware · 2026-09-12