HorizonMath Benchmark Released: GPT-5.4 Pro Discovers Novel Math Solutions

RexDouglass · x · 2026-08-03

Researchers introduced HorizonMath, a next-generation benchmark designed to measure AI capabilities in mathematical discovery. It consists of 101 unsolved problems where discovery is hard but verification is easy.

When testing frontier models like GPT-5.4 Pro, Claude 4.6 Opus, and Gemini 3.1 Pro, GPT-5.4 Pro stood out by finding two potentially novel solutions that beat existing baselines.

Original post →

More from Models

Models channel →