Best Agent Completes Only 61.68%: Alibaba's CommerceAgentBench Goes Live on GitHub
alifcoder · x · 2026-08-26
CommerceAgentBench is now open-source on GitHub (Accio-org/CommerceAgentBench, 1.2k stars). Its 107 tasks were distilled from Alibaba's real e-commerce data: 10M SME users → 1.6M full-length conversations → 200K commercial execution trajectories → 2,000 high-value workflows → 107 reproducible tasks, drawing on 27 years of Alibaba e-commerce experience. The leaderboard shows the best observed result completed just 66 of 107 tasks (61.68%) — a model can sound convincing in chat yet break when the job demands messy inputs, conflicting information, business rules, multiple tools, real execution, and a verifiable final state. The team invites contributors to expand its mock environments. The author's takeaway: agent benchmarks should ask less "what did the AI say?" and more "what did the AI actually do?"
Related event: Alibaba Open-Sources CommerceAgentBench for Real E-commerce Tasks(3 posts)→
More from coding & agent
- Developer shares experience of vibecoding a new agent orchestrator app — willcb · 2026-08-26
- NVIDIA claims up to 30× agentic throughput per MW on Vera Rubin — Crescitaly · 2026-08-26
- xAI launches Grok Bot: AI colleagues that use your tools — tetsuoai · 2026-08-26
- Dev builds Conduit: a browser-control extension for any agent, not just Claude — PumpkinNarrow6339 · 2026-08-26
- CommerceAgentBench Launch: Testing If AI Agents Can Complete Real E-commerce Tasks — alifcoder · 2026-08-26
- Running Qwen3.8-27B for local coding on 16GB VRAM: full setup guide — Due-Project-7507 · 2026-08-26