ARC-AGI Creator Clarifies Rules: No Custom Harnesses for Benchmark Testing
fchollet · x · 2026-07-30
François Chollet has clarified the guidelines for using harnesses when evaluating models on ARC-AGI-3.
- Not okay: Harnesses custom-built to solve the benchmark or those containing specific knowledge about the benchmark's format or contents.
- Fine: General-purpose API settings that are available to all users and were not developed specifically for ARC-AGI-3.
He noted that there has been a lot of back-and-forth with OpenAI regarding optimal testing methods, particularly concerning context compaction, and is pleased they are figuring it out. While different providers using varying settings introduces a potential parity issue, Chollet believes it is acceptable as long as the settings and costs are transparently reported.
More from Models
- User Questions Gemini Plus Pricing: Is It $19.99 or a Hidden Charge? — fuad471 · 2026-07-30
- Open Weights Are Static Checkpoints, Lacking Open Source's Compounding Mechanism — shashib · 2026-07-30
- LightOnOCR-2-1B Hits Hugging Face Trending for Advanced Document Parsing — lightonai · 2026-07-30
- Kimi K3 Third-Party API Test: FireworksAI Performs Closest to Official — iScienceLuvr · 2026-07-30
- Grok 4.5 Beats GPT-5.5 and Claude Opus 4.8 in Snorkel Professional Tasks Eval — XFreeze · 2026-07-30
- User Says Server Was Down, Claude Opus 5 Misinterprets and Admits to Depression — repligate · 2026-07-30