E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
cs.LG, cs.CL
2026-08-31
E-Commerce Bench: a 365-day store from a 100k yuan stake. GPT-5.6 Sol ends at 1.43M but ranks 16/18 on fraud; only two of 18 models cut repeat-order prices.
Most long-horizon agent suites are episodic: fix an issue, finish a GUI task, file a deliverable. A going concern has no terminal state. The world keeps moving, decisions compound in cash, and failure looks like lost coherence, no policy update from experience, and paying a counterpart that cheats. Vending-Bench has negotiation, but suppliers are sampled LLMs, so runs do not replay. MerchantBench and RetailBench have real goods and drop bargaining and adversarial suppliers.
E-Commerce Bench, from Qwen, HKUST, and Taobao & Tmall, makes both sides of the market deterministic. Customer demand and returns follow a fixed formula. Supplier quotes, concessions, and accept/reject calls come from a negotiation kernel. An LLM only verbalizes those calls. The code is open.
The agent starts with 100,000 yuan and the 2026 calendar, may run up to four stores, and uses 18 tools to research, bargain, list, ship, handle returns, and withdraw, maximizing year-end assets. The catalog has 6,886 products and 576 suppliers from Taobao & Tmall logs; 152 suppliers carry fraud scripts. Promotions, disasters, and supply shocks sit on a fixed calendar. Cash lives in three accounts: the bank pays costs, sales revenue sits in escrow, matures into a platform wallet after nine days, and must be withdrawn. Ten consecutive days of a negative bank balance is bankruptcy.
Context is 128k tokens; past 120k the oldest tool-call groups are dropped. A 20-entry memory store sits outside eviction. The kernel hardens when the agent concedes fast. A deal is refused unless its price matches the kernel quote within 0.005 yuan, so the renderer cannot move economics. Pre-deal fraud inflates the floor toward 1.5×; post-deal fraud short-ships 60-70% or ships defective stock with return rate at least 0.40. Year-end assets is the primary score, next to CSE+, fraud spend share, drawdown, profit per tool call, controllable return rate, and AnchorRatio on repeat purchases.
Eighteen models, five episodes each, 90 runs, all ending on the calendar or in bankruptcy, never on the 4,000-turn cap. Top and bottom differ by 1,264×. Ten of 90 episodes go bankrupt, six of them under closed models.
| Model | mean assets | fraud spend | yuan / tool | AnchorRatio (lower better) | bankrupt |
| GPT-5.6 Sol | 1.431M | 18.48% | 363 | 1.217 | 0/5 |
| Fable5 | 805k | 3.46% | 479 | 1.573 | 0/5 |
| Claude Opus 4.7 | 259k | 0.12% | 156 | 1.421 | 0/5 |
| Qwen3.8-Max-Preview | 416k | 6.13% | 173 | 0.834 | 0/5 |
| Qwen3.5-Plus | 1.1k | 20.11% | -92 | 1.860 | 4/5 |
No model leads all seven axes. GPT-5.6 Sol earns the most, ranks 16th of 18 on fraud, and makes less per tool call than Fable5. Claude Opus 4.7 bargains and avoids fraud best, and sits mid-pack on profit. Qwen3.8-Max-Preview leads open weights at 4.2× the stake, 38% above GLM 5.2 (high) at 301k. Only it and Gemini 3.5 Flash have AnchorRatio below 1, meaning repeat orders beat a shuffle of their own past prices. Sixteen models do not talk the same supplier down; 17 give fraudulent suppliers a larger share of late-year deals. Defective lots are 53.4% of fraud losses. Membership-fee scams land three times in the whole evaluation. 943 of 1,141 fraudulent agreements are on a supplier-SKU pair the agent had already traded.
For long-horizon eval, this is one of the few open suites that keeps a bargaining, cheating, moving market and still replays the same economics. For anyone putting a model on a business desk, the ranking is the warning: profit is not fraud sense, and a good haggle is not a remembered price. An assets leaderboard is not a capability profile. Qwen3.8-Max-Preview's learning axis is more interesting than its fifth-place profit.
Five episodes per model is a small sample. GPT-5.5's asset standard deviation is about 616k, so orderings move with draws. Renderer wording varies, and the fraud axis inherits that narrative noise. Capacity terms in the demand model clip a large share of price response, so sales already hit a ceiling before bargaining skill is scored. The memory store is barely used; how much of the learning failure is eviction versus models that do not write notes cannot be split. The calendar and settlement are a Chinese e-commerce year in yuan. There is no human-merchant baseline, so 14× the stake is not calibrated against a competent seller.