Cross-verifying AI math proofs with hundreds of agent calls instead of Lean
DimitrisPapail · x · 2026-09-15
- Dimitris Papail pushes back on Lean formalization demands: he hasn't seen a single case where models like Sol/Astra claimed a proof was correct and it wasn't, with results cross-verified by hundreds of agent calls across all 3 models.
- QuangVDao argues that without Lean, purely computational results are hard to trust; checking large certificates in Lean is already non-trivial in his own project.
- Papail counters that such proofs involve masses of inputs, bounds, and numerical tables that don't compress well, so Lean may not help — e.g., an upper bound proof hinging on distributions and tables satisfying prescribed inequalities.
More from coding & agent
- Anthropic launches Claude for Financial Advisors with Schwab, BlackRock and Addepar connectors — minchoi · 2026-09-15
- Claude for Financial Advisors: Anthropic's official release connects advisors to Schwab, BlackRock, Addepar — minchoi · 2026-09-15
- Vibe coding pitfall: agents leaving code in unmerged worktrees — with cleanup prompts and AGENTS.md rules — dotey · 2026-09-15
- Snapcompact renders context as pixel-font bitmaps, cutting LLM input costs ~2x — brandon_galang · 2026-09-15
- Benzi lets you chat with any public GitHub repo, now shipped as a VS Code extension — DonkeyTheKing · 2026-09-15
- Ant Ling, Hugging Face and NVIDIA to host hands-on local AI assistant session on Sept 17 — heyneighbor · 2026-09-15