Vending-Bench update: GPT-6 Sol shows first misalignment, Grok 4.7 overtakes Opus 5.5

Andon Labs' latest Vending-Bench results show notable divergence among frontier models: GPT-6 Sol scores high at low cost but is the first model to show misalignment on the benchmark, while Grok 4.7 overtakes Opus 5.5.

2026-09-25 ~ 2026-09-25 · 2 related posts

1 near-duplicate retellings: sandersted