DeepSeek vs Maka Benchmark: Inconsistent Baselines Skew Results

teortaxesTex · x · 2026-08-02

Addressing recent comparisons between DeepSeek and Maka model benchmarks, a developer pointed out statistical discrepancies in the evaluation metrics. DeepSeek scored 82.7% on Terminal Bench 2.1 using the official harness. In contrast, Maka's reported 85.3% pass rate was calculated only on a subset of 61 questions that completed without timeout. Because the denominators differ, this is not an apples-to-apples comparison, and directly comparing the scores can be misleading.

Original post →

More from Models

Models channel →