BrokenArXiv benchmark now tests last month's refuted arXiv conjectures, GPT-6 Astra on top

scaling01 · x · 2026-09-16

The latest BrokenArXiv and ArXivMath benchmarks now focus on arXiv conjectures refuted within the past month, and models are run inside a harness rather than via direct API calls. Performance remains strong, with GPT-6 Astra topping the board.

Original post →

More from Models

Models channel →