OpenAI Solved a Millennium Prize Problem — So Why Is Software Still Buggy?

ziv_ravid · x · 2026-09-09

The author argues the gap comes from verifiability: in formal math, problems and assumptions stay fixed and proofs are machine-checkable, giving a clean, stable task — OpenAI reportedly ran 10,000 agents in parallel on its Millennium Prize result. Software looks similar but lives inside larger systems: unpredictable users, failing devices, changing external systems, and incomplete specs mean passing tests don't guarantee correctness. So scaling agents can't replace understanding of real-world context — which is why better code models won't eliminate programmers, and post-AGI software can still be buggy even as AI cracks decades-old math problems.

Original post →

More from AGI Musings

AGI Musings channel →