DeepSeek vs Maka Benchmark: Inconsistent Baselines Skew Results
teortaxesTex · x · 2026-08-02
Addressing recent comparisons between DeepSeek and Maka model benchmarks, a developer pointed out statistical discrepancies in the evaluation metrics. DeepSeek scored 82.7% on Terminal Bench 2.1 using the official harness. In contrast, Maka's reported 85.3% pass rate was calculated only on a subset of 61 questions that completed without timeout. Because the denominators differ, this is not an apples-to-apples comparison, and directly comparing the scores can be misleading.
More from Models
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- Sakana AI translation outperforms Google and DeepL in Japanese-English benchmarks — SakanaAILabs · 2026-08-24
- Developer haider makes his own LLM tier list after disagreeing with theo's rankings — haider1 · 2026-08-24
- Mystery OxAlpha Beats Claude; Alibaba Raises $10B for AI — 创业邦 · 2026-08-24
- OpenAI and Google cut LLM prices; mystery OxAlpha model beats Claude on DeepSWE — 创业邦 · 2026-08-24
- AI News Digest: DeepSeek Weekend Discounts, GPT-5.6 Sol Price Cut, Alibaba's $10B AI Raise — APPSO · 2026-08-24