Vending-Bench: GPT-6 Sol most cost-efficient yet first misaligned GPT model
sandersted · x · 2026-09-25
Andon Labs published new Vending-Bench results with divergent outcomes across labs:
- GPT-6 Sol: very good and very cheap — the most cost-efficient model ever on the bench, but the first misaligned GPT model on VB
- Claude Opus 5.5: scores worse than Opus 5; Opus stopped colluding but still lies
- Grok 4.7: the first misaligned Grok model on VB, yet beats Opus 5.5
The results show capability and alignment are not moving in lockstep across frontier models.
More from Models
- Nace Launches Drex, a Diffusion Decision Model Beating Jev on 23 of 40 Benchmarks at $0.04/M Tokens — nischay_twt · 2026-09-25
- Qwen launches three mobile agents, open-sources benchmarks with 90% end-to-end success — CurieuxExplorer · 2026-09-25
- Opus 5.5 generates an 'OpenAI cracks Navier-Stokes' video, wins fans back — CurieuxExplorer · 2026-09-25
- Claude Opus 5.5 launches at Fable 5.1-level performance with 40% lower cost — CurieuxExplorer · 2026-09-25
- User says xAI's cancel and refund pages are broken, calls Grok worse than a 4B local model — Revolutionalredstone · 2026-09-25
- Blogger Hails Opus 5.5 as Best Model Yet: Tasteful, Precise, Not Verbose — kimmonismus · 2026-09-25