Qwen team's E-Commerce Bench runs 18 frontier models through a simulated year; no model dominates
dair_ai · x · 2026-09-02
The Qwen team released E-Commerce Bench, which runs an agent through a simulated 365-day year operating several online stores at once, scoring 18 frontier models across seven dimensions.
- GPT-5.6 Sol earns the most, growing a 100,000 stake into 1,431,425, but ranks 16th of 18 on fraud avoidance and trails Fable 5 on operational efficiency.
- Among open-weight models, Qwen3.8-Max-Preview leads at 416,252, 38% above GLM 5.2 (high), and shows the strongest learning over the horizon — progressively bargaining suppliers down across repeated orders.
- No single model dominates every dimension. Worth reading for anyone evaluating agents beyond a single session.
More from Models
- World Labs releases Atlas: A pixel-perfect camera controlled world model — viksit · 2026-09-02
- Fable 5.1 Science Benchmark: Autonomous success rate doubles to 53.6% — johnseach · 2026-09-02
- Anthropic releases Claude 5.1 with 25% lower costs and zero data retention — eugeneyan · 2026-09-02
- OpenAI previews Astra cybersecurity model reaching Critical threshold — OpenAI · 2026-09-02
- Observation: AI agents become succinct in voice mode, adapting to human listeners — joshwhiton · 2026-09-02
- Anthropic Accused of Retroactively Adding Safeguards to Older Opus Models — LordCoice · 2026-09-02