Minor Harness Setting Tweaks Radically Alter ARC-AGI Scores
teortaxesTex · x · 2026-08-01
Recent discussions around the ARC-AGI-3 benchmark reveal a crucial observation: minor changes to the execution framework's settings can significantly impact a model's final score.
Developers noted that simply making two adjustments to the Harness drastically improved Claude 3.5 Sonnet's performance. This raises concerns about the accuracy of current LLM evaluations and suggests that the true capabilities of models like DeepSeek V4 might be underestimated by existing testing frameworks.
More from coding & agent
- Cross-Repo Adversarial Review Workflow: Write with Claude, Review with Codex — dSebastien · 2026-08-01
- Tackling Alzheimer's: Crowdsourced Challenge Uses AI Agents to Map APOE4 Evidence — victormustar · 2026-08-01
- Formbar: Steering AI Video Generation via 3D Scenes from a Single Image — gorkem · 2026-08-01
- Claude + Thrixel Fully Automate Game Generation with Just a Handful of Prompts — RanaHanocka · 2026-08-01
- Exploring LEAN Agents: Building a Moat with Automated Formal Verification — teortaxesTex · 2026-08-01
- Non-Developer Seeks Alternatives for Deploying Personal AI Agents — Negative-Guard-4487 · 2026-08-01