MerchantBench Paper Reveals Long-Horizon Agent Performance Collapse

rohanpaul_ai · x · 2026-08-29

A new paper, MerchantBench, provides a brutal reality check on long-horizon AI capabilities. Testing eight leading models (including GPT-5.6 Sol and Claude Opus 4.8) in a simulated one-year e-commerce environment, the study found that even the best-performing setup, Qwen3.7-Max, ended with only 27.3% of the earnings of an average human participant. This exposes severe shortcomings in current agents regarding long-term coherence, delayed feedback, and complex decision chains. Citing Chamath, the author emphasizes that long-horizon tasks and complex problems remain unsolved, and current systems are far from reliable for long-term execution.

Original post →

More from coding & agent

coding & agent channel →