ARC-AGI 榜单引争议:评测框架比模型本身更能决定分数
thursdai_pod · x · 2026-08-10
The recent controversy surrounding the ARC-AGI leaderboard isn't actually about the models themselves, but rather the evaluation harness.
The ThursdAI podcast invited guests to break down the benchmark issues, emphasizing that how the testing framework is constructed dictates the final scores significantly.
「研究」频道最新
- 传闻:中国开源模型靠逆向提取推理链实现蒸馏突破 — jxmnop · 2026-08-10
- 从 GPT-2 到 Kimi3:万字长文解析大模型架构演进史 — iamrobotbear · 2026-08-10
- TBSM新方法实现单步图像生成,20B文生图模型推理加速 — 机器之心 · 2026-08-10
- 实测 SQLite 压缩历史记录:1000 次修订文本从 20MB 压缩至 80KB — Simon Willison · 2026-08-10
- 个人项目 KLQ:无需训练的大模型量化新法,超越 SpinQuant — Federal-Setting-3014 · 2026-08-10
- AI 也要打工赚工资:开源项目 ClawWork 测试模型真实盈亏 — dr_cintas · 2026-08-10