Dispute: Claiming 100% on benchmark without private eval is invalid
JFPuget · x · 2026-08-31
JFPuget responded to criticism about their model's score on the ARC benchmark, clarifying they only tested on public evaluation games and noting the lack of tools for semi-private data evaluation. François Chollet countered that claiming a score on a benchmark without private evaluation is invalid and criticized the decision not to open source the solution.
Related event: Chollet responds to ARC-AGI benchmark scoring dispute(4 posts)→
More from Models
- Developer Switches from Opus 5 to Codex Citing Better Performance — sirbayes · 2026-08-31
- Dots3-Note Weights Open: Addressing the Benchmark-Reality Gap — CodeByPoonam · 2026-08-31
- Tiel-Coder-35B-A3B Trends on HF with Speculative Decoding & MTP — peculiar-ragdoll · 2026-08-31
- Google Releases Gemini Omni 1.1 Flash, Updating Its Fast Multimodal Model for Developers — thione · 2026-08-31
- DeepSeek launches low-cost vision model; Anthropic previews hardware control protocol for agents — thione · 2026-08-31
- Qwen and GLM release new MoE models focusing on low cost and high performance — thione · 2026-08-31