FULL STORY
ARC-AGI-3 Scores Spark Controversy and Reflection
High scores on the ARC-AGI-3 benchmark sparked cheating and overfitting controversies. Amid developer skepticism, officials clarified rules and open-sourced the codebase, prompting industry reflection.
2026-07-25 ~ 2026-08-01 · 9 episodes · 64 posts
Episode 1 · Opus 5's High ARC-AGI-3 Score Sparks Cheating and Overfitting Controversy (2026-07-25, 11 posts)
Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, outperforming peers by 3 times. However, this high score immediately triggered widespread allegations of cheating and overfitting within the AI community, with multiple users and industry insiders questioning its true capabilities and sparking reflections on the validity of evaluation systems and their ties to commercial interests.
Confirmed
It is confirmed that Opus 5's specific score is 30.2%, and this result has caused a severe crisis of confidence in the industry. Reddit user @sdnr8 questioned the opacity of closed-source models, suggesting the high score might not come from "pure model" capabilities but from external proxy frameworks like loop or harness. User @VraserX directly accused Anthropic of "cheating," claiming they specifically trained for the benchmark's puzzle patterns, converting visual reasoning into explicit algebraic problems. Furthermore, a person claiming to be from a frontier AI lab told @flowersslop that the score "looks fake," implying either excessive benchmark gaming or a severe underestimation of Anthropic's lead by competitors. Chris Szegedy (@ChrSzegedy) and @DrSingularity explicitly stated that ARC-AGI has little to do with AGI, though the latter acknowledged that AI is indeed getting stronger. Gary Marcus also emphasized that scoring high on benchmarks does not equate to approaching AGI, arguing that the score improvement is more likely due to targeted optimization rather than the true generalization of abstract reasoning capabilities.
Unconfirmed
Whether Opus 5 actually used external proxy frameworks, or whether Anthropic indeed conducted targeted specialized training, currently remains at the stage of community speculation and accusations, with no substantial technical evidence provided for the relevant claims.
Why it matters
This controversy affects more than just Anthropic's reputation; it reflects the industry's deep-seated concerns about the validity of AI evaluation systems. On one hand, as relayed by @morqon, rapidly saturating evaluations lose their discriminative power, with some believing the benchmark was nearly "solved" on day one. On the other hand, discussions forwarded by @burnytech indicate a growing default assumption that ARC-AGI progress mainly comes from targeted RL environments, prompting calls for ARC-AGI-4 to reduce public demos to prevent model overfitting. @JasonBotteril also pointed out that naming a benchmark directly after AGI is almost destined to induce labs to optimize for it specifically. In addition, @rbhar90 emphasized that the "strong reasoning ability" narrative of frontier labs is deeply bound to huge commercial interests, advocating for more open, verifiable benchmarks (such as the self-built Witness test set), and noting that Opus 5's improvement on ARC-AGI-3 did not transfer to other test sets. @inductionheads also noted that if a model is indeed trained in an RL environment similar to ARC-AGI, its performance cannot prove "generalization." This indicates that existing public benchmarks are facing severe methodological challenges.
- Anthropic’s ARC-AGI-3 lead is being called meaningless as the benchmark saturates — morqon · 2026-07-25
- User Accuses Anthropic of Gaming ARC-AGI-3 by Training Specifically on Benchmark Patterns — VraserX · 2026-07-25
- Reddit post questions whether Opus 5’s ARC-AGI-3 score came from a looped harness — sdnr8 · 2026-07-25
- Frontier lab rumor says Opus 5 ARC-AGI 3 score looks fake — flowersslop · 2026-07-25
- ARC-AGI-4 should stay private, after Opus 5 scored 3× the next-best model on ARC-AGI-3 — burny_tech · 2026-07-25
- Opus 5 reaches 30.2% on ARC-AGI 3 as critics question the benchmark — ChrSzegedy · 2026-07-25
- A reply says ARC-AGI is not AGI, even as better AIs keep pushing the field forward — Dr_Singularity · 2026-07-25
- Open benchmarking challenge says Opus 5’s ARC-AGI-3 leap doesn’t carry over — rbhar90 · 2026-07-25
- Anthropic’s Opus 5 ARC-AGI Debate Reopens the Generalization Question — inductionheads · 2026-07-25
- Gary Marcus says benchmark gains are not the same as reaching AGI — GaryMarcus · 2026-07-27
- A benchmark named after AGI may have guaranteed labs would optimize for it — JasonBotterill · 2026-07-27
Episode 2 · Deep Dive into Opus 5 Hidden Reasoning and ARC-AGI Score (2026-07-25, 2 posts)
Technical analyses explore Opus 5's hidden reasoning mechanisms and agent-like behaviors, while subsequent clarifications detail its 30.16 ARC-AGI-3 score using the RHAS evaluation and compute trends.
- Opus 5, hidden-rule inference, and J-space point to a new agent stack — imjustnewatai · 2026-07-25
- Clarifying Opus 5's ARC-AGI Score Details and Compute Scaling Math — imjustnewatai · 2026-07-25
Episode 3 · Anthropic's Benchmark Scores Spark Community Trust Crisis (2026-07-25, 2 posts)
Discussions around Anthropic's benchmark scores have sparked a trust crisis within the community. Despite decent results that reportedly outperform Opus 4.8, users tend to question the benchmark's reliability, with high ECI metrics further amplifying score discrepancies and confusion.
- A benchmark joke says Anthropic doing badly is the fastest way to lose trust in it — teortaxesTex · 2026-07-25
- Anthropic benchmark score looks fine, but high ECI makes small gaps look bigger — scaling01 · 2026-07-25
Episode 4 · Gary Marcus Says ARC-AGI Name Is Misleading (2026-07-27, 2 posts)
Gary Marcus and others argue that the name "ARC-AGI" is misleading, as it makes people mistakenly believe the benchmark directly tests for AGI, a concept that remains vaguely defined and highly debated.
- ARC-AGI’s name may overstate what the benchmark can really tell us about AGI — tedgreenwald · 2026-07-27
- Gary Marcus says ARC-AGI’s name makes people think it tests AGI itself — GaryMarcus · 2026-07-27
Episode 5 · Human Baselines Missing in AI Evaluations, Highlighting Human-AI Synergy (2026-07-29, 4 posts)
Experts warn that complex AI benchmarks are losing crucial human baselines for comparison. To address this, researchers are advocating for new evaluation standards like Stanford's CollabSkill framework, which focuses on measuring the collaborative performance of humans and AI agents in real-world tasks.
- Humans-plus-AI need their own evals, not just model benchmarks — paraschopra · 2026-07-29
- Stanford Paper Introduces CollabSkill: Evaluating Human-Agent Collaboration — paraschopra · 2026-07-30
- Frontier AI Benchmarks Lose Meaning Without Human Baselines — emollick · 2026-07-30
- Ethan Mollick: Frontier AI Benchmarks Are Losing Human Baselines — emollick · 2026-07-31
Episode 6 · Optimized Memory Settings Triple GPT-5.6's Score on ARC-AGI-3 (2026-07-30, 27 posts)
OpenAI officially confirmed that by enabling "reasoning retention" and "context compaction" in the Responses API, GPT-5.6 Sol's score on the ARC-AGI-3 public test set surged from 13.3% to 38.3% (a 188% increase), reaching SOTA levels while reducing token consumption to 1/6th of the original. This reveals that the memory management mechanism within the test harness, rather than pure model reasoning ability, is the key determinant of performance in complex long-horizon tasks.
Confirmed
- GPT-5.6 Sol initially performed poorly on the ARC-AGI-3 benchmark, an issue first highlighted in tests by developer @sandersted.
- By enabling two specific API settings used internally by ChatGPT and Codex, the model's score increased roughly threefold (from 13.3% to 38.3%), alongside a sixfold improvement in token efficiency (reducing consumption to 1/6th).
- The two key settings are reasoning retention and context compaction, which allow the model to remember previous thought processes and operate continuously across multiple context windows.
- OpenAI engineer @ilanbigio noted that standard test harnesses discard the model's reasoning process after each step and lose early operation records when the context fills up, limiting the model's evaluated performance in isolated environments.
Why it matters
- This discovery challenges the current industry-standard isolated evaluation methods. Developers like @Angaisb pointed out that complex tasks can be easily solved by optimizing the test harness, meaning that merely comparing raw benchmark scores does not truly represent the actual intelligence of different AI systems.
- The deep synergy between product frameworks (like ChatGPT and Codex's internal mechanisms) and model capabilities has a decisive impact on complex task performance, indicating that the performance bottleneck often lies in context management rather than the model itself.
- GPT-5.6 ARC-AGI-3 scores surge 3x with specific API settings enabled — sandersted · 2026-07-30
- GPT-5.6 ARC-AGI-3 Score Jumps 3x With Specific API Settings — sandersted · 2026-07-30
- Test: Enabling Two API Settings Boosts GPT-5.6 ARC-AGI-3 Score by 3x — sandersted · 2026-07-30
- GPT-5.6 ARC-AGI-3 Score Triples When Enabling Thought Memory — sandersted · 2026-07-30
- OpenAI Reveals Compaction Boosts GPT-5.6 ARC-AGI-3 Score to 38% — ilanbigio · 2026-07-30
- OpenAI Reveals Compaction Boosts GPT-5.6 ARC-AGI-3 Score to 38% — ilanbigio · 2026-07-30
- Optimized Harness Achieves SoTA on ARC-AGI-3, Questioning Benchmark Validity — Angaisb_ · 2026-07-30
- OpenAI Tripled Its ARC-AGI-3 Scores by Enabling Just Two Settings — ObiWanCanownme · 2026-07-30
- OpenAI: Optimized API Settings Triple GPT-5.6 Sol's Score on ARC-AGI-3 — OpenAI · 2026-07-30
- OpenAI Optimizes ARC-AGI-3 Harness, Boosts Score by 188% with 6x Fewer Tokens — OpenAI · 2026-07-30
- GPT-5.6 Sol Achieves SOTA on ARC-AGI-3 via Context Compaction — daniel_mac8 · 2026-07-30
- Retaining Reasoning Across Turns Boosts GPT-5.6 Sol on ARC-AGI-3 by ~3x — charliermarsh · 2026-07-30
- GPT-5.6 Sol Reasoning Details: Lack of Memory Forces Re-learning Every Step — charliermarsh · 2026-07-30
- Harness Matters: GPT-5.6 Hits SoTA on ARC-AGI-3 with Two Setting Tweaks — TheZachMueller · 2026-07-30
- OpenAI Triples ARC-AGI-3 Score with 6x Fewer Tokens via API Tweaks — tw_killian · 2026-07-30
- GPT Performance Jumps 188% with 6x Fewer Tokens via Context Compaction — soumitrashukla9 · 2026-07-30
- Two API Settings Triple GPT-5.6's ARC-AGI-3 Score While Cutting Tokens — gabrielchua · 2026-07-30
- GPT-5.6 Sol Achieves SOTA on ARC-AGI-3 with Context Compaction Tweaks — soumitrashukla9 · 2026-07-30
- OpenAI Triples GPT-5.6 Score on ARC-AGI-3 by Tweaking API Settings — GregKamradt · 2026-07-30
- Two API Tweaks Boost OpenAI Model's ARC-AGI-3 Score and Slash Compute — xiaohu · 2026-07-30
Episode 7 · ARC-AGI 3 Evaluation Mechanism Under Fire from Developers (2026-07-30, 9 posts)
Recent developers have strongly questioned the fairness and evaluation mechanism of the ARC-AGI 3 benchmark, arguing that the current test framework design, scoring rules, and even testing motives are severely distorted and cannot truly measure general AI capabilities.
Confirmed
- Framework limits memory: Developers @Glittering-Neck-2505 and @teortaxesTex point out that the current unified test framework deliberately prevents reasoning agents from maintaining context across operations, forcing models to constantly "forget." This practice, aimed at cross-provider standardization (using the same completion-style endpoint and not passing back reasoning logs), limits agent long-term memory and fails to effectively measure true model efficiency.
- Evaluation framework significantly affects scores: Developers @ImaginaryDinner2710 and @ilanbigio (via @dkundel) note that model scores heavily depend on the test harness. For ARC-AGI, an OpenAI model scored only 13% due to improper invocation, jumping to 38% after correction. Additionally, just two adjustments to the harness settings led to a significant improvement for Claude 3.5 Sonnet. Properly tracking reasoning and using context compression are crucial for achieving true high scores.
- Scoring mechanism criticized: Developer @mgostIH found that current harnesses have saturated ARC-AGI-3; a small amount of reinforcement learning (RL) can greatly boost scores. He also strongly criticizes the benchmark's use of squaring the percentage of solved problems, calling it baseless and "myth-making."
- Alleged malicious tuning: Developer @iruletheworldmo harshly criticizes ARC-AGI as a catalyst for public attention, suggesting the testers may have maliciously lowered model scores to create the illusion that the benchmark is extremely hard to saturate.
- Unfair base model testing: ML Street Talk (via @burnytech) argues that modern AI is essentially a hybrid neuro-symbolic system, and a single-context reasoning window is insufficient for adaptation, so directly testing base models on ARC-AGI3 is unfair.
Unconfirmed
- Whether testers subjectively "maliciously" tuned parameters to create the illusion of extreme difficulty (currently only a harsh accusation by @iruletheworldmo).
- Whether some models used a different harness in internal preview tests than in public evaluations (question raised by @LiangSong850509).
- Whether the industry can reach consensus on establishing and adopting a unified agent framework for testing different models (currently only a suggestion from ML Street Talk).
Why it matters
- Affects benchmark credibility: If the test mechanism can be easily gamed with a small amount of RL and saturates, the scoring formula is controversial, and there is even potential for artificially lowering scores, the benchmark's authority will be greatly diminished.
- Evaluation must reflect real usage: Developers emphasize that in real development scenarios, models always run with specific agent frameworks. When designing cross-platform standardized tests, one should not sacrifice core agent capabilities (like long-term memory). The industry urgently needs to abandon "bare" benchmarks.
- How Test Harness Design Impacts ARC-AGI-3 Scores — dkundel · 2026-07-30
- Reddit Debate: Is ARC-AGI 3 an Intentionally Dishonest Measure of AGI? — Glittering-Neck-2505 · 2026-07-30
- Agent Benchmark Harness Criticized: Gimping Long-Term Memory Skews Efficiency Tests — teortaxesTex · 2026-07-30
- Dev Critiques ARC-AGI-3: Already Saturated, Quadratic Scoring Inflates Egos — mgostIH · 2026-07-30
- ML Street Talk: ARC-AGI3 Should Be Benchmarked with a Unified Agentic Harness — burny_tech · 2026-07-30
- Harness Errors Cause 25% Score Swings, Exposing Flaws in Raw Benchmarks — Imaginary_Dinner2710 · 2026-07-31
- ARC-AGI Benchmark Accused of Bad Faith Rigging for Publicity — iruletheworldmo · 2026-07-31
- Agent Performance Shaped by Both Models and Execution Frameworks — LiangSong850509 · 2026-07-31
- Minor Harness Setting Tweaks Radically Alter ARC-AGI Scores — teortaxesTex · 2026-08-01
Episode 8 · Claude Opus ARC-AGI Score Questioned Over API Flaw (2026-07-30, 2 posts)
A developer argued that Claude Opus's high ARC-AGI score was inflated by an API flaw, but community members maintain that its generalization capabilities still outperform GPT-5 under the same testing framework.
- Netizen Questions: If Harness is Identical, Opus Generalizes Better Than GPT-5 — umike_njsf · 2026-07-30
- ARC-AGI Evaluation Dispute: Claude Opus Score Questioned Over API Implementation Flaw — steipete · 2026-07-30
Episode 9 · ARC-AGI-3 Benchmark Rules Clarified and Official Code Released (2026-07-30, 5 posts)
Recent evaluation methods for large models on the ARC-AGI-3 benchmark have sparked controversy. In response, ARC-AGI founder François Chollet and official member Greg Kamradt have clarified the evaluation rules and API usage, and open-sourced the official testing codebase to ensure fairness and transparency across model evaluations.
Confirmed
- Custom tools banned: François Chollet explicitly prohibits using custom tools developed specifically for this benchmark or systems containing knowledge of the test format and content.
- Fair testing mechanism: Addressing the controversy over Claude Opus, Greg Kamradt clarified that the official testing used the exact same "sliding window" convention for both Anthropic and OpenAI models.
- Open-source evaluation tool: Greg Kamradt shared the official ARC Prize open-source codebase arc-agi-3-benchmarking, which provides complete environment setup and running guides, supporting configuration of various API keys to evaluate large models on complex abstract reasoning tasks.
- API field clarification: Greg Kamradt corrected misconceptions about the optional reasoning field in the API, emphasizing it is not for extracting the model's internal chain-of-thought (CoT) but only for recording and associating external outputs.
Why it matters
A unified and open evaluation standard helps eliminate suspicions of "cheating" among model providers during benchmarking. Explicitly banning custom tools and publishing test code provides developers with a reproducible evaluation environment, further solidifying ARC-AGI's authority as a benchmark for frontier model intelligence.
- Greg Kamradt Responds to Evaluation Dispute: Same Rolling Window Used for Opus and OpenAI — GregKamradt · 2026-07-30
- ARC-AGI-3 Official Benchmarking Repo Goes Open Source — GregKamradt · 2026-07-30
- ARC-AGI-3 Official Benchmarking Repository Open-Sourced — GregKamradt · 2026-07-30
- ARC-AGI-3 API clarification: 'reasoning' field logs model output, not private CoTs — GregKamradt · 2026-07-30
- ARC-AGI Creator Clarifies Rules: No Custom Harnesses for Benchmark Testing — fchollet · 2026-07-30