Qwen's E-CommerceBench makes 18 models run an online shop for a year
千问大模型 · wechat · 2026-09-03
The Qwen team, with Taotian Group, released E-CommerceBench: agents start with ¥100K and run an online shop for 365 days against a simulated marketplace built on desensitized Taobao/Tmall data (6,886 products, 576 suppliers, 152 fraudulent). Vendor negotiation is driven by a deterministic kernel with LLM-only dialogue rendering, ensuring reproducibility and blocking jailbreak-based price hacking. Across 18 models × 5 runs, GPT-5.6Sol topped total assets at 14.31x but ranked 16th on fraud detection; 10 runs went bankrupt from January overstocking. Only Qwen3.8-Max-Preview showed genuine long-horizon learning (AnchorRatio 0.88), actually driving purchase prices down over the year. No model ranks well on all seven dimensions, showing single-metric scores hide real weaknesses.
Related event: Qwen Releases E-Commerce Bench: No Model Masters Year-Long Store Operation(4 posts)→
More from Research
- Cohere Labs releases ATE dataset with ~700K tools from public MCP servers — Cohere_Labs · 2026-09-03
- IFM open-sources K2 Horizon models from 0.9B to 375B with training code and data recipes — testingcatalog · 2026-09-03
- Recursive LM author clarifies: what experiments prove vs. what the work is about — CShorten30 · 2026-09-03
- Qwen-RobotManip: alignment before scale for robotic manipulation foundation models — rsasaki0109 · 2026-09-03
- Late-interaction doc embeddings shrink 255x to 6KB per page, keeping 95%+ accuracy — lateinteraction · 2026-09-03
- NeoMME-Retriever-260M tops sub-800M models on ViDoRe v3 with 0.523 nDCG@10 — lateinteraction · 2026-09-03