RealReplicaBench: All AI Agents Fail E-commerce Test, Top Score 56.1

APPSO · wechat · 2026-08-17

APPSO reports on RealReplicaBench, an e-commerce AI agent benchmark by Alibaba's AccioWork. Across 107 real business tasks, all models failed, with top scorer Claude Opus 5 at 56.1. The benchmark emphasizes task completion, requiring results usable by next steps, by replicating real environments and verifying final states. The team plans to update tasks for model evaluation and optimization.

Related event: AI models fail real-world work benchmarks as e-commerce agents fall short(2 posts)→

Original post →

More from Research

Research channel →