ARC-AGI Creator Clarifies Rules: No Custom Harnesses for Benchmark Testing

fchollet · x · 2026-07-30

François Chollet has clarified the guidelines for using harnesses when evaluating models on ARC-AGI-3.

He noted that there has been a lot of back-and-forth with OpenAI regarding optimal testing methods, particularly concerning context compaction, and is pleased they are figuring it out. While different providers using varying settings introduces a potential parity issue, Chollet believes it is acceptable as long as the settings and costs are transparently reported.

Original post →

More from Models

Models channel →