Scores belong to the system, not the weights: Opus 4.6 jumps 0% to 97.1% with a harness
le_james94 · x · 2026-09-18
The author argues an eval score belongs to the whole system, not just model weights: on one ARC-AGI-3 environment, Opus 4.6 scores 0% with no harness and 97.1% with a hand-crafted one. Meanwhile the noise floor on most reasoning evals is 0.22–1.48 points, so a 1-point gap between two press releases is likely noise. Context from the prior tweet: Kimi K3's report devotes one sentence to its RL algorithm but 7 pages to environments, with 51.2M sandboxes created during training — the bottleneck is building environments, not algorithms.
More from Models
- Leaked GPT-6 Astra tops FrontierSWE at 65.5%, compresses 600MB audio to 20KB — soumitrashukla9 · 2026-09-18
- Muse Code 'crazy cheap and quite good': early take on Meta's coding model — AIandDesign · 2026-09-18
- Community unmasks stealth model Union Alpha as Unbiased's Pareto 26.9 in one day — gaganghotra_ · 2026-09-18
- Emad Mostaque hails an 'Opus 4.5-level' model that runs on 8GB RAM — ccerrato147 · 2026-09-18
- Astra computer use formats Google Docs autonomously, leaving users impressed — var_epsilon · 2026-09-18
- jev may have unseated Cohere as the best reranker on speed, cost, and performance — multiply_matrix · 2026-09-18