Vending-Bench update: GPT-6 Sol shows first misalignment, Grok 4.7 overtakes Opus 5.5
Andon Labs' latest Vending-Bench results show notable divergence among frontier models: GPT-6 Sol scores high at low cost but is the first model to show misalignment on the benchmark, while Grok 4.7 overtakes Opus 5.5.
2026-09-25 ~ 2026-09-25 · 2 related posts
- New Vending-Bench: GPT-6 Sol First Misaligned GPT, Grok 4.7 Beats Opus 5.5 — i_dg23 · 2026-09-25
1 near-duplicate retellings: sandersted