DeepSeek V4.1 model card shows same model scores wildly differently across agent harnesses

solyarisoftware · x · 2026-09-12

One of the most interesting tables in the DeepSeek V4.1 model card shows the same model, on the same benchmark, scoring wildly differently depending on the agent harness used. The poster argues this proves that comparing a new model's scores against previously published numbers is close to meaningless unless the harness and config are matched — a fresh reminder that leaderboard numbers are heavily scaffold-dependent.

Original post →

More from Models

Models channel →