MerchantBench: 365-Day E-Commerce Simulation Tests LLM Agents' Long-Term Coherence
rohanpaul_ai · x · 2026-10-02
A team unveils MerchantBench, a 365-day order-level simulation for seller-side e-commerce built on 98,843 real product records with 26 interaction tools, benchmarking LLM agents' Long-Term Coherence — sustaining purposeful behavior over extended horizons while adapting to accumulated evidence. It couples promptly observable upstream supplier events with delayed downstream order outcomes, requiring agents to track order lifecycles and revisit earlier decisions across sourcing, listing/pricing, and cash-flow management. Eight LLMs were evaluated under two agent frameworks in 48 runs of 365 simulated days, with models generally struggling at long horizons.
More from coding & agent
- Dev recreates a DOOM-like game with GPT, nailing the original's gory feel — DeryaTR_ · 2026-10-02
- Process-mining agents found 20 steps and 7 loops in a workflow documented as 7 steps — vasuman · 2026-10-02
- Geoffrey Huntley: Forget reading code—your verification properties are all that matters — kieranklaassen · 2026-10-02
- Claude Skills explained: why prewritten PDF scripts beat pasting prompts every time — lxfater · 2026-10-02
- Eye.Art Polyphemus: a chat-first MCP for image generation and reference-based edits — axiomofaxiom · 2026-10-02
- Dev uses GitHub Copilot as a project lead: files issues, lets the agent do everything else — 0xkarasy · 2026-10-02