A math researcher’s AI workflow review ranks models on differential geometry
BLUECOW009 · x · 2026-07-21
John Ennis wrote up his experience using AI for mathematics, with the quoted article focusing on what worked and what failed in real math workflows.
The attached ranking image compares several models on differential geometry tasks:
- 5.6 Sol was best at proof architecture and final repair, but could overfit to a narrower statement than intended.
- Claude Fable 5 handled deep derivations and long rewrites well, but sometimes relied on an elegant yet unproven global claim.
- Sakana Fugu Ultra was strong at hostile review and checking hidden issues, but was slow and overly literal.
- Kimi K3 was noted for proposing unusual original mechanisms.
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22