Qwen and DeepSeek Benchmarks Reportedly Use Claude Code Harness

goddamnit_1 · reddit · 2026-08-18

A user observed that benchmarks for open models like Qwen 2.5 and DeepSeek utilize Claude Code as the evaluation harness, instead of their native or open-source alternatives like OpenCode or Hermes Agent. The post questions whether this choice is driven solely by performance gains and asks the community about their real-world experiences with different harnesses.

Original post →

More from Models

Models channel →