Minor Harness Setting Tweaks Radically Alter ARC-AGI Scores

teortaxesTex · x · 2026-08-01

Recent discussions around the ARC-AGI-3 benchmark reveal a crucial observation: minor changes to the execution framework's settings can significantly impact a model's final score.

Developers noted that simply making two adjustments to the Harness drastically improved Claude 3.5 Sonnet's performance. This raises concerns about the accuracy of current LLM evaluations and suggests that the true capabilities of models like DeepSeek V4 might be underestimated by existing testing frameworks.

Related event: ARC-AGI 3 Evaluation Mechanism Questioned: Framework Limits and Scoring Rules Distorted(8 posts)→

Original post →

More from coding & agent

coding & agent channel →