Luna Model Matches Heavyweights in Web Agent Benchmarks at 1/20th the Cost

kohjingyu · x · 2026-07-31

The Luna model achieves strong scores of 51% and 55.4% on the Odysseys and MyPCBench computer-use benchmarks. It performs similarly to the heaviest flagship models while being over 20x cheaper to run.

The Odysseys benchmark itself evaluates long-horizon web tasks consisting of 200 real-world browsing scenarios. Even the strongest frontier models achieve only a 44.5% perfect task success rate, indicating substantial room for improvement in long-horizon web agents.

Original post →

More from Models

Models channel →