Verifying GPT Pro: Zero Math Errors, But Disastrous Exposition
josh_wills · x · 2026-08-04
Mathematician Dimitris Papail shared his observations while verifying specific mathematical theorems using GPT Pro:
- High Accuracy: Zero mathematical mistakes found so far. When GPT Pro asserts something is correct, he trusts it more than his own Lean verifier.
- Disastrous Exposition: The model's proofs are exhausting to read. It uses a complicated tree of variable renamings (e.g., renaming norm ratios to Greek letters), forcing the reader to track confusing changes.
- Random Structure: The ordering of technical lemmas is chaotic, lacking the progressive logical structure expected in human mathematical exposition.
Despite the exhausting experience, he marvels at the model's ability to destroy problems in under an hour that have stymied stronger human minds for much longer.
More from Models
- llama.cpp patches boost DeepSeek-V4-Flash-0731 from 3.26 to 25.91 tok/s — dyn___ · 2026-08-04
- Claude Fable 5 tops MirrorCode with a 64% solve rate, far ahead of GPT-5.6 Sol — davidad · 2026-08-04
- Qwen3.8-Max ranks No. 2 in Vision Arena, 13 points behind Claude Fable 5 — arena · 2026-08-04
- Polymarket puts an 82% chance on a new Google Gemini Pro model within two weeks — Polymarket · 2026-08-04
- Users look for uncensored VLMs that can caption explicit images accurately — TekeshiX · 2026-08-04
- Claude usage meter may be overstating quota after just two turns this morning — dreamwieber · 2026-08-04