MerchantBench tests LLM agents over 365 days of simulated e-commerce operations
dair_ai · x · 2026-08-04
MerchantBench proposes a 365-day simulation for evaluating agentic e-commerce performance under long-term coherence constraints.
- Built on 98,843 real product records
- Exposes agents to 26 tools for sourcing, listing, pricing, cash-flow management, and delayed feedback
- Scored by cumulative net assets, so mistakes compound over time instead of averaging out
- The authors ran 8 LLMs across 2 agent frameworks for 48 full-year simulations
- Best result still reached only 27.3% of the human baseline in mean final net assets
Related event: Alibaba Launches MerchantBench for E-commerce Agents(3 posts)→
More from coding & agent
- Building a (deliberately unsafe) restricted shell MCP server for local coding agents — ag789 · 2026-09-22
- Rumor: OpenAI, Anthropic, and Cognition to launch personal agent platforms within a month — altryne · 2026-09-22
- One test decides if you need an AI agent: if you can write the steps down, you don't — Virtual_Hair_1987 · 2026-09-22
- Teknium merges fix for Hermes Agent Desktop failing to resolve model-provider plugins — Teknium · 2026-09-22
- 90s multi-agent systems offer a lesson: model memory as belief, not centralized truth — nptacek · 2026-09-22
- Grok 4.7 jumps to 46.3% on CursorBench and 64% on EEBench, keeping the same $2/$6 per million token pricing — FinanceYF5 · 2026-09-22