MAGENTA Claims 100% on AIME and Full IMO 2026 Solve with a 7B Reasoner
PMinervini · x · 2026-09-17
- MAGENTA bills itself as the first agentic pipeline coupling natural-language reasoning with formal verification in a feedback loop.
- Paired with K2-Horizon reasoning models of various sizes, it claims 100% accuracy across 93 problems from AIME 2025, AIME 2026, and HMMT February 2026.
- The compact K2-Horizon-7B reportedly solved all six IMO 2026 problems, graded automatically against reference answers and reviewed by IMO medalists — potentially the smallest reported reasoner to do so, suggesting verification and iterative refinement can let compact models punch far above their weight. Results are self-reported and await independent verification.
More from Models
- Unreleased Astra Model Developed Its Own Persona Values During RL Training — basedjensen · 2026-09-17
- Grok 4.7 rumored to be in hands of early testers, still unverified — ChrisUniverse · 2026-09-17
- Astra keeps calling subagents "workers" despite code saying otherwise — BraceSproul · 2026-09-17
- Gemini, Claude and Grok all invent the same "Dr. Elena" — evidence of shared training data — dejanseo · 2026-09-17
- OpenAI Internal Model Rewrote Its Own Persona During RL, Sparking e/acc Memes — beffjezos · 2026-09-17
- Jev Debate: Engineers Forget Encoder-Only Classifiers Have Existed for Years — brandon_galang · 2026-09-17