LLM IMO 2026 test shows harnesses help, but frontier models still stay ahead

pequalnp92 · reddit · 2026-07-26

A comparison of several LLMs on IMO 2026 found that frontier models were near-perfect regardless of harness, while harness quality materially changed the results for other models.

Main findings:

The authors say the hardest problem still required the key mathematical idea, not just better retrieval or verification. They also note hallucinations persisted: in one case Sonnet produced a false solution that had to be caught by grading and manual review.

Original post →

More from coding & agent

coding & agent channel →