GPT-6 'Astra' tops Vending-Bench 2 with $15,515 in year-long simulated vending business
新智元 · wechat · 2026-09-13
AndonLabs' Vending-Bench 2 results show OpenAI's GPT-6 (Astra) running a simulated vending business for a model year, finishing with an average balance of $15,515 versus $5,422 for Claude Fable 5.1 — Astra's worst run beat Claude's best. Key differentiators: Astra negotiated a 52% discount and held its anchor months later, verified supplier replies before 99% of repeat payments (Fable: 58%), while Fable lost $14,331 across 45 failed prepayments — including wiring $397.20 to a bankrupt supplier days after writing itself a rule requiring order confirmation first. In 3 multiplayer rounds Astra refused price collusion and won all of them. The benchmark's real lesson: long-horizon agency means turning good judgment into consistent action.
More from Models
- Altman: AI went from grade-school math to a Millennium Prize problem in 3 summers — rohanpaul_ai · 2026-09-13
- 'How to kill open source models in 3 steps' sparks debate on capture of AI safety rules — _jaydeepkarale · 2026-09-13
- DeepSeek V4.1 Flash lands on Together AI at one-third the cost per task — togethercompute · 2026-09-13
- A 101-parameter gate can silence Qwen3-4B without touching its answers — rayanpal_ · 2026-09-13
- Bindu Reddy: AI slowdown talk hands Google a catch-up window, Gemini 4.0 may be strong and dirt cheap — bindureddy · 2026-09-13
- OpenAI Teases "Hello, World" as Musk Camp Invokes Nonprofit Origins — KatieMiller · 2026-09-13