New study says human TSP solutions are near-optimal but still systematically human
xuanalogue · x · 2026-07-29
A new arXiv preprint studies how humans produce near-optimal, human-like solutions to combinatorial optimization problems, using the Euclidean TSP as the main case.
The authors sampled a broad set of TSP instances, collected human tours, and compared them with Pointer Network policies trained under several objectives: reinforcement learning, supervised learning from optimal tours, supervised learning from human tours, and RL fine-tuning after optimal pretraining. They find that human tours are not identical to optimal tours, but occupy a near-optimal geometric basin with systematic human-specific deviations.
The best explanation for human tours was not direct imitation of optimal tours, but a model pretrained on optimal tours, fine-tuned with RL, and decoded with Best-of-N sampling.
More from Research
- NSF launches 4-year PhD program as predictions say every enterprise may soon run an AI lab — annbordetsky · 2026-07-29
- New paper says AI-written books are flooding Amazon and crowding out sales — TuhinChakr · 2026-07-29
- SenseTime open-sources SenseNova-Vision, a unified model for detection to 3D reconstruction — socialwithaayan · 2026-07-29
- Schmidhuber says AlphaFold overlooked earlier protein-structure prediction work — SchmidhuberAI · 2026-07-29
- New Paper Uses Inverse RL to Extract Auditable Alignment Rewards from Demos — nagpalchirag · 2026-07-29
- AI Tackles Undeciphered Languages Like Linear A and Etruscan — Ars Technica AI · 2026-07-29