Theorem Generation Agent Rankings on VeriBench
sanmikoyejo · x · 2026-07-12
A forwarded post showcases benchmark results from VeriBench, comparing the performance of theorem generation agents.
The results show: DSPy ReAct 0.615 > Trace++ 0.588 > Trace+ 0.477 > baseline 0.470. The author emphasizes that these conclusions only represent calibrated evidence on the validation set and do not imply universal reliability; a locked hold-out test set will be rolled out next.
More from Research
- giffmana skeptical: found training env already contaminated, eval protections unlikely to hold — giffmana · 2026-09-11
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- Fast ViT shows strong ImageNet results; scaling runs needed next — ducha_aiki · 2026-09-11
- Loss Functions Are Scientific Assumptions: MSE Implies Gaussian Noise, Cross-Entropy Implies Bernoulli — bravo_abad · 2026-09-11
- SymKit MCP: 44 tools for AI agents to verify symbolic derivations — Foreign-Specific-604 · 2026-09-11
- Researchers: LLMs under pressure invent new languages unreadable to humans — mikeflache · 2026-09-11