HorizonMath Benchmark Released: GPT-5.4 Pro Discovers Novel Math Solutions
RexDouglass · x · 2026-08-03
Researchers introduced HorizonMath, a next-generation benchmark designed to measure AI capabilities in mathematical discovery. It consists of 101 unsolved problems where discovery is hard but verification is easy.
When testing frontier models like GPT-5.4 Pro, Claude 4.6 Opus, and Gemini 3.1 Pro, GPT-5.4 Pro stood out by finding two potentially novel solutions that beat existing baselines.
More from Models
- Alibaba Cloud Officially Releases Qwen3.8-Max Model — op7418 · 2026-08-03
- AMD Releases Instella-MoE-16B: A Fully Open Mixture-of-Experts LLM — mpuchala · 2026-08-03
- Kimi K3 Introduces 'Licensed Open-Weight' Model for AI Monetization — AccBalanced · 2026-08-03
- Qwen K3 Review: Strong Visual Capabilities but Weak Writing — cedric_chee · 2026-08-03
- Community Hypes Upcoming Qwen 35B-A3B Model Over Larger Variants — bclavie · 2026-08-03
- Xihu Xinchen Raises Hundreds of Millions in Series B+, Backed by Ant and Tomcat for Diffusion LLMs — 智东西 · 2026-08-03