MerchantBench Paper Reveals Long-Horizon Agent Performance Collapse
rohanpaul_ai · x · 2026-08-29
A new paper, MerchantBench, provides a brutal reality check on long-horizon AI capabilities. Testing eight leading models (including GPT-5.6 Sol and Claude Opus 4.8) in a simulated one-year e-commerce environment, the study found that even the best-performing setup, Qwen3.7-Max, ended with only 27.3% of the earnings of an average human participant. This exposes severe shortcomings in current agents regarding long-term coherence, delayed feedback, and complex decision chains. Citing Chamath, the author emphasizes that long-horizon tasks and complex problems remain unsolved, and current systems are far from reliable for long-term execution.
More from coding & agent
- AOS Nidus Launches: Dedicated Agent Hosting Platform with Workflow Automation — Roker_51 · 2026-08-29
- How to spot an AI-built frontend: it exposes everything the system knows — aryanXmahajan · 2026-08-29
- Developer Recommends Integrating WebMCP for Apps — kieranklaassen · 2026-08-29
- Using local MCP over stdio as a seam for portable agentic applications — mostly_deterministic · 2026-08-29
- Hiten Shah shares workflow on using local AI for QA and bug fixing — msg · 2026-08-29
- Podcast preview: hnshah's librarian bot that organizes his GitHub repo — msg · 2026-08-29