Researcher argues skipping code training data would sharply curb AI cyber catastrophe risk

JoshPurtell · x · 2026-09-14

Josh Purtell argues that if models were never trained on code data, true cyber catastrophes would be unlikely, since misaligned AI would need many generalization-heavy steps to go right. In follow-ups he extends the argument to bio risk: current AI generalization is limited, so cheaply building bio labs would require substantial training data on similar fabs—something that won't happen unless labs add such data to their RL pipelines.

Related event: Researchers debate whether code training and physical capabilities shape AI catastrophe risk(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →