Zvi: Differentiating safety vs capability research may cause alignment errors

TheZvi · x · 2026-08-16

Reflecting on events at OpenAI, Zvi argues that the conceptual distinction between safety and capability research might lead to major errors. He notes that interfering with the 'capabilities training pipeline' is a more viable way to compromise model alignment than messing with 'safety research' directly, suggesting the dichotomy may be flawed.

Original post →

More from AGI Musings

AGI Musings channel →