New study says human TSP solutions are near-optimal but still systematically human

xuanalogue · x · 2026-07-29

A new arXiv preprint studies how humans produce near-optimal, human-like solutions to combinatorial optimization problems, using the Euclidean TSP as the main case.

The authors sampled a broad set of TSP instances, collected human tours, and compared them with Pointer Network policies trained under several objectives: reinforcement learning, supervised learning from optimal tours, supervised learning from human tours, and RL fine-tuning after optimal pretraining. They find that human tours are not identical to optimal tours, but occupy a near-optimal geometric basin with systematic human-specific deviations.

The best explanation for human tours was not direct imitation of optimal tours, but a model pretrained on optimal tours, fine-tuned with RL, and decoded with Best-of-N sampling.

Original post →

More from Research

Research channel →