Alibaba's MerchantBench Reveals E-Commerce Agents Attain Only 27% of Human Performance
alibabagroup · hf · 2026-08-05
Alibaba introduced MerchantBench, a novel benchmark designed to evaluate Long-Term Coherence in LLM agents for e-commerce operations.
Background & Mechanism
- While existing benchmarks focus on bounded tasks, real-world deployments require agents to preserve purposeful behavior across extended horizons.
- The benchmark features a 365-day order-level simulation grounded in 98,843 real e-commerce product records.
- Agents are equipped with 26 tools to handle interdependent decisions like product sourcing, pricing, and cash-flow management.
Results
- The team evaluated eight LLMs across 48 runs.
- Results reveal a substantial gap between the latest LLMs and human participants.
- The best LLM configuration attained only 27.3% of the mean final net assets achieved by humans.
Related event: Alibaba Launches MerchantBench for E-commerce Agents(3 posts)→
More from coding & agent
- Flight Intelligence MCP: Flight Search & Comparison Agent Tool via Google Flights — modelcontextprotocol · 2026-08-06
- Blacksmith MCP: Let Claude Query CI/CD Analytics and Workflow Logs Directly — modelcontextprotocol · 2026-08-06
- Gemini Image Gen Combined with GPT Coding Easily Creates AI Visual Heroes — RichardsonDx · 2026-08-06
- LangChain Founder Releases Open Source Agent Starter Kit — hwchase17 · 2026-08-06
- Cursor's Grok 4.5 High Drastically Increases Token Usage Per Turn — NickPassig · 2026-08-06
- Muse Code Matches Codex and Claude in Real-World Coding Test — alexandr_wang · 2026-08-06