MerchantBench: Benchmarking LLM Agents for Long-Term E-Commerce Coherence
_akhaliq · x · 2026-08-06
MerchantBench is a novel benchmark designed to evaluate the long-term coherence of LLM agents in e-commerce operations.
- Environment: Simulates 365-day, order-level e-commerce operations, grounded in 98,843 real product records.
- Tools: Equips agents with 26 merchant tools for product sourcing, dynamic pricing, order tracking, and cash-flow management.
- Challenges: Tests agent resilience against supplier disruptions and delayed outcomes like refunds, negative reviews, and penalties.
- Paper and code are open source.
Related event: Alibaba Launches MerchantBench for E-commerce Agents(3 posts)→
More from coding & agent
- Developer gets Codex remote control on phone: "My life is over" — gowthami_s · 2026-09-22
- Vibe-coded DuckDB extension lands paying customers within five days — josh_wills · 2026-09-22
- GPT-6 Astra learns to play Fallout 2 from screenshots and keyboard/mouse inputs alone — Vjeux · 2026-09-22
- Google's RRSI regularizes recursive self-improvement of agent harnesses, +4.7 OOD points — google · 2026-09-22
- How Do You Run User-Facing Agent Harnesses in Production? Reddit Thread Asks — YamSpiritual1964 · 2026-09-22
- 53 harness bugs: an engineering team's field notes on building eval harnesses — DrDatta_AIIMS · 2026-09-22