LLMs Fall Short in Writing Security Proofs
matthew_d_green · x · 2026-07-14
Matthew Green recently tried using ChatGPT and Fable to write new security proofs (not correctness proofs) targeting EasyCrypt or Lean, but the results were "less than ideal."
This post offers highly specific practical feedback:
- LLMs are not yet reliable for formal security proofs, at least in the author's use case.
- The issue isn't a complete inability to write them, but a noticeable gap from being practically usable.
- The author also remains open to the possibility that their prompting might have been suboptimal.
Related event: LLMs Struggle with Complex Cryptographic Security Proofs(2 posts)→
More from Research
- Statistical theory paper studies how fast signatures learn in path regression — chaumian · 2026-07-21
- PROWL uses a world model to keep Minecraft agents exploring after failures — nathanbenaich · 2026-07-21
- LeRobot v0.6.0 adds end-to-end 3D depth training data for robots — RemiCadene · 2026-07-21
- Multiagent v2 playbook calls for 64 agents, diverse proof routes and adversarial checks — danshipper · 2026-07-21
- HarmonicMath says Lean autonomously solved eight previously studied open problems — MarioKrenn6240 · 2026-07-21
- SeeSE3 finds 3D structure emerging in frozen vision features and camera-pose alignment — ducha_aiki · 2026-07-21