8 Lean formalizations in OpenAI's repo called "broken"; Opus 5.5 weighs in on Comparator critiques
ctjlewis · x · 2026-10-09
Timeroot reports that 8 Lean formalizations in OpenAI's repo are "broken" — trivially provable with wrong proofs. Elliot Glazer then used Claude Opus 5.5 to compare this critique with Robert George's earlier one, concluding that Alex is right that OpenAI botched the Comparator setup multiple times, though incidentally — each would likely have passed a legitimate Comparator run.
Related event: Eight OpenAI Lean Formalizations Found Trivially Provable with False Proofs(3 posts)→
More from Research
- MOSS@COLM 2026 spotlights small-scale LLM research with talks by Danqi Chen and Omar Khattab — sewon__min · 2026-10-10
- Data firms need frontier-level training infra to prove data value, says Proximal — AccBalanced · 2026-10-10
- AI agent misses $14M in lease clauses: Surge AI's GDP.xlsx benchmarks agents on real spreadsheets — echen · 2026-10-10
- Calico open-sources Cerberus DNA model in PyTorch: 786kb input, 8361 human tracks — anshulkundaje · 2026-10-10
- New preprint uses state space models for long-sequence genomic prediction, open-sources Cerberus PyTorch models — anshulkundaje · 2026-10-10
- MaCVi maritime computer vision workshop lands at WACV 2027 with 3D sonar and underwater benchmarks — HildeKuehne · 2026-10-10