Greg Kamradt Responds to Evaluation Dispute: Same Rolling Window Used for Opus and OpenAI
GregKamradt · x · 2026-07-30
Addressing the recent controversy over Claude Opus allegedly cheating in the ARC-AGI evaluation, Greg Kamradt clarified the methodology. He stated that they used the exact same "rolling window" convention for both Anthropic and OpenAI models to ensure fairness in the test.
Related event: ARC-AGI Open Sources Eval Code Amid Testing Dispute(2 posts)→
More from Models
- Google Gemini Robotics Set for Major 2.0 Update — CyberRobooo · 2026-07-30
- Cisco's Open-Source Security Model Antares-3B Nears GPT-5.5 at Fraction of Cost — joshua_saxe · 2026-07-30
- Grok Outpaces Claude Opus 5 in App Coding Speed Test — yunta_tsai · 2026-07-30
- User Reports Claude $200 Max Plan Usage Limits Stealth Nerfed — ATTlKA · 2026-07-30
- Dev Slams Open AI Releases: Demands Affordable API Endpoints Like $5 Kimi-K3 — arthurcolle · 2026-07-30
- Anthropic's New Opus Models Accused of 'Laziness' and Slashing Workloads in Automation — 歸藏的AI工具箱 · 2026-07-30