MerchantBench tests LLM agents over 365 days of simulated e-commerce operations
dair_ai · x · 2026-08-04
MerchantBench proposes a 365-day simulation for evaluating agentic e-commerce performance under long-term coherence constraints.
- Built on 98,843 real product records
- Exposes agents to 26 tools for sourcing, listing, pricing, cash-flow management, and delayed feedback
- Scored by cumulative net assets, so mistakes compound over time instead of averaging out
- The authors ran 8 LLMs across 2 agent frameworks for 48 full-year simulations
- Best result still reached only 27.3% of the human baseline in mean final net assets
More from coding & agent
- Enkstein open-sources a local control plane for Codex, Claude, Gemini, and Ollama — wcoreiron · 2026-08-04
- Gemini Spark can connect MCP servers, but every tool still needs repeated approval — tristanbob · 2026-08-04
- Agent apps should expose a skill surface, then crystallize it into code — burny_tech · 2026-08-04
- Use subagents to switch models inside one session — dotey · 2026-08-04
- xAI updates Grok Build with Grok 4.5, skills, MCP, and plan mode — elonmusk · 2026-08-04
- RL on custom search harnesses may beat the “one big model” idea — shangbinfeng · 2026-08-04