Qwen and DeepSeek Benchmarks Reportedly Use Claude Code Harness
goddamnit_1 · reddit · 2026-08-18
A user observed that benchmarks for open models like Qwen 2.5 and DeepSeek utilize Claude Code as the evaluation harness, instead of their native or open-source alternatives like OpenCode or Hermes Agent. The post questions whether this choice is driven solely by performance gains and asks the community about their real-world experiences with different harnesses.
More from Models
- User Review: GPT Luna outperforms Claude Sonnet in tests — oyacaro · 2026-08-18
- User surprised by how good GPT Luna is: smart and pleasant to use — oyacaro · 2026-08-18
- Solar Pro 4 Available for Free for One Week on Nous Portal — NousResearch · 2026-08-18
- Hermes Just Works: A Brief Endorsement — NERDDISCO · 2026-08-18
- Running AI Agents on CPU: Gemma 4 26B vs Qwen 3.6 35B? — Kahvana · 2026-08-18
- Qwen3.8-27B benchmarks and SGLang high-throughput serving guide — Sam Witteveen · 2026-08-18