CommerceAgentBench released: Qwen leads open-weight models
Alibaba_Qwen · x · 2026-09-01
Alibaba's Accio team open-sourced CommerceAgentBench, a benchmark for real commerce operations. It runs agents in high-fidelity, stateful replicas of real online services.
Key Findings:
- The best overall completion rate is only 62%, highlighting the difficulty of execution.
- Qwen3.8-Max delivered the strongest performance among open-weight models evaluated.
- The benchmark focuses on execution in real commercial workflows, not just answering questions.
Related event: Alibaba's Accio Open-Sources CommerceAgentBench for E-commerce Agents(3 posts)→
More from coding & agent
- VibeKit MCP Server Manages Deployments, Logs, and Headless Coding — modelcontextprotocol · 2026-09-01
- Distributed.systems发布可审计的Agent基础设施 — arthurcolle · 2026-09-01
- How to Stop Context Window Bottlenecks in Data-Heavy MCP Servers — JuicerSocial · 2026-09-01
- Grok Bots Turn LLM Citations Into an SEO Loop for AI Search Ranking — rohanpaul_ai · 2026-09-01
- Search configuration impacts agent accuracy 40x more than model choice — RichardSocher · 2026-09-01
- Built a LoL Classic Wiki & Build Planner using Claude for data and logic — Shortykane · 2026-09-01