Harness Errors Cause 25% Score Swings, Exposing Flaws in Raw Benchmarks

Imaginary_Dinner2710 · reddit · 2026-07-31

The author points out that current LLM benchmarks are highly dependent on the test harness. On the ARC-AGI benchmark, an OpenAI model scored only 13% due to improper API calls, but jumped to 38% when corrected.

In real-world development, models are always used within specific agent frameworks (like Claude Code). Testing raw models without these harnesses is meaningless, fails to reflect actual capabilities, and misleads developers. The industry needs to move away from this unrepresentative evaluation method.

Related event: ARC-AGI-3 Evaluation Framework Criticized for Memory Limits and Flawed Scoring(6 posts)→

Original post →

More from Models

Models channel →