Scores belong to the system, not the weights: Opus 4.6 jumps 0% to 97.1% with a harness

le_james94 · x · 2026-09-18

The author argues an eval score belongs to the whole system, not just model weights: on one ARC-AGI-3 environment, Opus 4.6 scores 0% with no harness and 97.1% with a hand-crafted one. Meanwhile the noise floor on most reasoning evals is 0.22–1.48 points, so a 1-point gap between two press releases is likely noise. Context from the prior tweet: Kimi K3's report devotes one sentence to its RL algorithm but 7 pages to environments, with 51.2M sandboxes created during training — the bottleneck is building environments, not algorithms.

Original post →

More from Models

Models channel →