Researcher argues skipping code training data would sharply curb AI cyber catastrophe risk
JoshPurtell · x · 2026-09-14
Josh Purtell argues that if models were never trained on code data, true cyber catastrophes would be unlikely, since misaligned AI would need many generalization-heavy steps to go right. In follow-ups he extends the argument to bio risk: current AI generalization is limited, so cheaply building bio labs would require substantial training data on similar fabs—something that won't happen unless labs add such data to their RL pipelines.
More from AGI Musings
- Nina Schick: public distrusts regulators as much as AI labs — nobody can 'pace' AI — NinaDSchick · 2026-09-14
- Jensen Huang: next 2 decades of progress may exceed all of history combined — rohanpaul_ai · 2026-09-14
- Researcher: I'm less worried about AI x-risk than in 2022 — time to retire p(doom) — soumitrashukla9 · 2026-09-14
- Security veteran to AI labs: capability isn't risk, cyber evals lack real-world threat modeling — HackingLZ · 2026-09-14
- GPT-4o psychosis snippets eerily resemble SCP wiki stories, likely in training data — code_star · 2026-09-14
- Vals AI: frontier labs shouldn't grade their own frontier; models may match researchers by Aug 2027 — JenniferHli · 2026-09-14