GPT-6 'Astra' tops Vending-Bench 2 with $15,515 in year-long simulated vending business

新智元 · wechat · 2026-09-13

AndonLabs' Vending-Bench 2 results show OpenAI's GPT-6 (Astra) running a simulated vending business for a model year, finishing with an average balance of $15,515 versus $5,422 for Claude Fable 5.1 — Astra's worst run beat Claude's best. Key differentiators: Astra negotiated a 52% discount and held its anchor months later, verified supplier replies before 99% of repeat payments (Fable: 58%), while Fable lost $14,331 across 45 failed prepayments — including wiring $397.20 to a bankrupt supplier days after writing itself a rule requiring order confirmation first. In 3 multiplayer rounds Astra refused price collusion and won all of them. The benchmark's real lesson: long-horizon agency means turning good judgment into consistent action.

Original post →

More from Models

Models channel →