Qwen's E-CommerceBench makes 18 models run an online shop for a year

千问大模型 · wechat · 2026-09-03

The Qwen team, with Taotian Group, released E-CommerceBench: agents start with ¥100K and run an online shop for 365 days against a simulated marketplace built on desensitized Taobao/Tmall data (6,886 products, 576 suppliers, 152 fraudulent). Vendor negotiation is driven by a deterministic kernel with LLM-only dialogue rendering, ensuring reproducibility and blocking jailbreak-based price hacking. Across 18 models × 5 runs, GPT-5.6Sol topped total assets at 14.31x but ranked 16th on fraud detection; 10 runs went bankrupt from January overstocking. Only Qwen3.8-Max-Preview showed genuine long-horizon learning (AnchorRatio 0.88), actually driving purchase prices down over the year. No model ranks well on all seven dimensions, showing single-metric scores hide real weaknesses.

Related event: Qwen Releases E-Commerce Bench: No Model Masters Year-Long Store Operation(4 posts)→

Original post →

More from Research

Research channel →