Alibaba's MerchantBench Reveals E-Commerce Agents Attain Only 27% of Human Performance
alibabagroup · hf · 2026-08-05
Alibaba introduced MerchantBench, a novel benchmark designed to evaluate Long-Term Coherence in LLM agents for e-commerce operations.
Background & Mechanism
- While existing benchmarks focus on bounded tasks, real-world deployments require agents to preserve purposeful behavior across extended horizons.
- The benchmark features a 365-day order-level simulation grounded in 98,843 real e-commerce product records.
- Agents are equipped with 26 tools to handle interdependent decisions like product sourcing, pricing, and cash-flow management.
Results
- The team evaluated eight LLMs across 48 runs.
- Results reveal a substantial gap between the latest LLMs and human participants.
- The best LLM configuration attained only 27.3% of the mean final net assets achieved by humans.
Related event: Alibaba Launches MerchantBench for E-commerce Agents(3 posts)→
More from coding & agent
- Developer gets Codex remote control on phone: "My life is over" — gowthami_s · 2026-09-22
- Vibe-coded DuckDB extension lands paying customers within five days — josh_wills · 2026-09-22
- GPT-6 Astra learns to play Fallout 2 from screenshots and keyboard/mouse inputs alone — Vjeux · 2026-09-22
- Google's RRSI regularizes recursive self-improvement of agent harnesses, +4.7 OOD points — google · 2026-09-22
- How Do You Run User-Facing Agent Harnesses in Production? Reddit Thread Asks — YamSpiritual1964 · 2026-09-22
- 53 harness bugs: an engineering team's field notes on building eval harnesses — DrDatta_AIIMS · 2026-09-22