Opus 5 appears to improve on ARC-AGI 1 and 2, and may rely on algebraic puzzle solving

herbiebradley · x · 2026-07-25

Herbie Bradley says the Opus 5 system card shows clear gains on ARC-AGI 1 and 2 compared with 4.7/4.8.

He also notes that the system-card description of an ARC-AGI-3-like task sounds like the model learns to turn a visual puzzle into explicit algebra, which he sees as a very trainable strategy. Based on that, he predicts the improvement likely came from a focused research project on these puzzle types, possibly using a data vendor under an exclusivity agreement.

Original post →

More from Models

Models channel →