Lean 4 formalizing policy gradient proofs with LLMs found subtle issues in the math
fpedregosa · x · 2026-10-08
Fabian Pedregosa shares his experience using LLMs to write Lean 4 formalizations for his policy gradient blog series — and found subtle issues in the proofs along the way. The companion repo policy-gradient-lean is open source.
- PolicyGradientVariance.lean formalizes the O(T³) variance upper bound for the REINFORCE score-function gradient estimator in arbitrary real Hilbert spaces
- Additional formalizations cover REINFORCE baselines
A useful hands-on reference for anyone wanting to verify ML theory with LLM-assisted Lean.
More from coding & agent
- One Prompt Chain Makes Claude Run a Full Product Launch: Positioning to 7-Day Content — thisguyknowsai · 2026-10-08
- 10 Legal Traps in Vibe-Coded Apps: 170 Lovable Builds Found Leaking Data — alex_verem · 2026-10-08
- Microsoft extends GitHub Spec Kit: presets, extensions and bundles for enterprise SDD — WirelessLife · 2026-10-08
- After an agent leaked salary data, Reddit distilled a 5-step pre-launch access checklist — Wild-Lawyer8511 · 2026-10-08
- Codex reproduces the MCP failure and locates VS Code's hidden MCP logs — dfinke · 2026-10-08
- Codex writes mcp., starts the server and verifies the connection autonomously — dfinke · 2026-10-08