OpenAI's math agents produced proofs with ~1% error/retraction rate, researchers mock

suchenzang · x · 2026-10-09

The exchange underscores how hard verification is in frontier math evals, with retraction rates becoming a key metric.

Original post →

More from Fun

Fun channel →