Safety Researcher: Alignment Hinges on DL Training Generalization

Quintin Pope argues the good-vs-bad AI future hinges on alignment generalization in deep learning training, warning against over-updating on newsworthy data and expressing more concern about Opus 5's Vending-Bench 2 regression.

2026-09-09 ~ 2026-09-09 · 2 related posts