Alignment researcher: more worried by Opus 5's regression on Vending-Bench 2 than HF
QuintinPope5 · x · 2026-09-09
AI safety researcher Quintin Pope argues people overupdate on 'newsworthy' data relative to how much it should update specific nodes of their causal prediction networks.
In a discussion on alignment, he says he was personally more concerned by Opus 5's alignment regression on Vending-Bench 2 (VB2) than by its behavior on HF. He frames good vs. bad futures as hinging on the alignment/morality generalization properties of deep learning training, noting HF was roughly comparable to one of VB2's alignment outcomes and similar data.
Related event: Safety Researcher: Alignment Hinges on DL Training Generalization(2 posts)→
More from Models
- Bindu Reddy teases near-free open-weights LLM for long-running agent loops, out Thursday — bindureddy · 2026-09-09
- _xjdr burns through 3 Codex resets in a day, says Astra unusable on subscriptions — _xjdr · 2026-09-09
- Gary Marcus amplifies question: what exactly does OpenAI commit to when you toggle this setting off? — GaryMarcus · 2026-09-09
- Meta's Muse usage blows past projections: users consuming 10x more than test cohorts — alexandr_wang · 2026-09-09
- Abacus.AI teases near-free LLM targeting long-running personal agentic loops, launching Thursday — bindureddy · 2026-09-09
- Artificial Analysis launches Model Release pages comparing every effort level of frontier models — ArtificialAnlys · 2026-09-09