Dispute: Claiming 100% on benchmark without private eval is invalid

JFPuget · x · 2026-08-31

JFPuget responded to criticism about their model's score on the ARC benchmark, clarifying they only tested on public evaluation games and noting the lack of tools for semi-private data evaluation. François Chollet countered that claiming a score on a benchmark without private evaluation is invalid and criticized the decision not to open source the solution.

Related event: Chollet responds to ARC-AGI benchmark scoring dispute(4 posts)→

Original post →

More from Models

Models channel →