RL Agent Hits Provably Optimal Quadrilateral Meshes on 90 of 96 Test Domains, Crushing Gmsh
Arjun Narayanan · hf · 2026-09-29
- Researchers trained an RL agent to build quadrilateral block decompositions of planar domains that reach the provable floor on vertex irregularity set by the discrete Gauss-Bonnet identity (called par).
- The agent edits a half-edge mesh structure directly; its policy network's convolutions follow mesh connectivity, so it generalizes unchanged to domains far larger than training.
- Training combines behavior cloning on trivially constructible optimal meshes (walked backward into demonstrations) with PPO, solving the sparse-reward exploration barrier.
- Results: on 96 held-out domains it completed an all-quad mesh on every one, usable on 95.7% on average, and provably optimal on 90. Gmsh's strongest configuration completed 51, usable 38, optimal 0 — and never produced a more regular mesh even with 3-14x more elements. On 64 domains twice training size it completed all, usable on 62, with median excess over par below 1 vs Gmsh's 39.
More from Research
- HCOMP 2026 study measures proof burden in public bounty listings on RentAHuman — windx0303 · 2026-09-29
- Statistician mocks frequentists for patching one flaw and ignoring the rest — RexDouglass · 2026-09-29
- Pre-registration is just year-old priors, statistician argues in Bayesian-frequentist spat — RexDouglass · 2026-09-29
- 300M-param image model trained for €5000 claims SD 1.5-level performance — incorporo · 2026-09-29
- Cheap verifiers match costly ones in LLM post-training, saving up to 99.7% of grading cost — iScienceLuvr · 2026-09-29
- Telescopic LM trains one model valid at every depth, cutting quality-budget area 43% — iScienceLuvr · 2026-09-29