BrokenArXiv benchmark now tests last month's refuted arXiv conjectures, GPT-6 Astra on top
scaling01 · x · 2026-09-16
The latest BrokenArXiv and ArXivMath benchmarks now focus on arXiv conjectures refuted within the past month, and models are run inside a harness rather than via direct API calls. Performance remains strong, with GPT-6 Astra topping the board.
More from Models
- Users report ChatGPT sessions getting muddled, answering questions from other chats — koltregaskes · 2026-09-16
- OpenAI reportedly prepping Codex Replay to run and compare historical task threads in parallel — testingcatalog · 2026-09-16
- Mystery stealth model Union Alpha hits OpenRouter: free, 256K context, agentic focus — gaganghotra_ · 2026-09-16
- Developer builds demo hours after getting Jev access, drawing researcher banter — suchenzang · 2026-09-16
- TabPFN-3.5 tops Kaggle's Otto competition out of the box, but experts call the benchmark flawed — RichmanRonald · 2026-09-16
- Bindu Reddy teases Opus 5.2 in testing, Grok 4.8 weeks away, OpenAI's Astra+ in testing — bindureddy · 2026-09-16