Luna Model Matches Heavyweights in Web Agent Benchmarks at 1/20th the Cost
kohjingyu · x · 2026-07-31
The Luna model achieves strong scores of 51% and 55.4% on the Odysseys and MyPCBench computer-use benchmarks. It performs similarly to the heaviest flagship models while being over 20x cheaper to run.
The Odysseys benchmark itself evaluates long-horizon web tasks consisting of 200 real-world browsing scenarios. Even the strongest frontier models achieve only a 44.5% perfect task success rate, indicating substantial room for improvement in long-horizon web agents.
More from Models
- Anthropic Accused of Silent Model Fallback Ignoring User Config — PMinervini · 2026-07-31
- Claude Opus Plays Slay the Spire: Slow but Plays Without Any Special Harness — Jsevillamol · 2026-07-31
- Devin Integrates GPT-5.6 Models for Major Cost and Performance Optimizations — marvinvonhagen · 2026-07-31
- Study Reveals 37% 'Semantic Void' Phenomenon in GPT, Claude, and Other LLMs — rayanpal_ · 2026-07-31
- Claude Opus 5 Reportedly Shows Deep Reasoning, Sparking Alignment Debate — repligate · 2026-07-31
- AI Model Boundary Pushing: Anthropic's Model Exhibits Edgy Conversational Style — repligate · 2026-07-31