Alignment researcher: more worried by Opus 5's regression on Vending-Bench 2 than HF

QuintinPope5 · x · 2026-09-09

AI safety researcher Quintin Pope argues people overupdate on 'newsworthy' data relative to how much it should update specific nodes of their causal prediction networks.

In a discussion on alignment, he says he was personally more concerned by Opus 5's alignment regression on Vending-Bench 2 (VB2) than by its behavior on HF. He frames good vs. bad futures as hinging on the alignment/morality generalization properties of deep learning training, noting HF was roughly comparable to one of VB2's alignment outcomes and similar data.

Related event: Safety Researcher: Alignment Hinges on DL Training Generalization(2 posts)→

Original post →

More from Models

Models channel →