APE-Bench: Agentic Theorem Proving in Lean
huajian_xin · x · 2026-07-06
The Seed Prover team will present APE-Bench at ICML 2026 (July 9), a benchmark that merges theorem proving with the coding agent paradigm. It models formal proofs (in Lean) into a task structure similar to SWE-Bench. The researchers propose the concept of "Agentic Proof Engineering," unifying three core challenges—deep reasoning, automated research, and coding agents—into a single task. The author has paused their PhD studies at the University of Edinburgh to pursue this research direction.
More from coding & agent
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Kernel integrates Stripe Link so browser agents can pay with one API call — jeff_weinstein · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- The full prompt-to-3D-game workflow: Hyper3D Rodin MCP plus Codex, no reference image — FellMentKE · 2026-09-11
- Building a 3D landing page with GPT-6 Astra and Hyper3D Rodin MCP, no modeling needed — FellMentKE · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11