ARC-AGI-3's non-standard scoring under fire as GPT-6 hits ~100%
peterwildeford · x · 2026-09-14
- Peter Wildeford (ex-ARC researcher) remarked that ARC-AGI-3's scoring system is "way sillier" than he realized.
- Sterling Croxton notes ARC-AGI-3 uses a totally non-standard way of calculating scores, making a hypothetical jump to 100% by models like GPT-6 the most likely outcome regardless of true capability gains.
- Critics argue ARC should never have adopted this methodology, and it has now "blown up in their faces."
- Takeaway: ARC-AGI-3 scores may not be comparable to previous benchmarks and should be interpreted cautiously.
More from Models
- Ollama's jmorgan: small models now handle most conversational and reasoning use cases — ollama · 2026-09-14
- GLM Flash 5.3 keeps slipping into Chinese mid-conversation, user reports — DevDminGod · 2026-09-14
- AI Claims First Millennium Prize: Navier-Stokes Falls, OpenAI Model Leaps in Math — TheZvi · 2026-09-14
- If embeddings are so powerful, why is retrieval their only mainstream use? — ProposalOrganic1043 · 2026-09-14
- An OOD test worth watching: asking Gemini 4.1 Pro to paint surrealism with pure Python code — teortaxesTex · 2026-09-14
- Leaked GPT-6 Sol Output Impresses, But OpenAI Reportedly Has Stronger Bell Internally — VraserX · 2026-09-14