Six years on, wunderwuzzi's ML attack series on model backdoors and Assume Breach still holds up
wunderwuzzi23 · x · 2026-10-08
Security researcher wunderwuzzi (Embrace The Red) revisits his 2020 Machine Learning Attack Series, written to bridge the gap between ML safety research and security practice. His core stance: adopt an Assume Breach mindset and ask "how would we even know it happened, and what do we do about it?"
The series uses a self-built image classifier (Husky AI) to cover ML threat modeling and hands-on attacks:
- Attack surface: pipeline attacks, brute-forcing and perturbations to force misclassifications, image scaling attacks, GAN-generated deceptive samples, adversarial examples with Microsoft Counterfit
- Model supply chain: stealing model files, backdooring Keras/Pickle model files (and detecting it), backdooring Jupyter Notebook files
- Real-world cases: CVE-2020-16977 VS Code Python extension RCE, signing binaries to bypass malware models
- Governance: a companion "Assume Bias" post argues AI systems will inevitably err, so designers must ship mitigation, detection, and response (including deprecation) mechanisms up front
Notably, he observes these analyses remain highly applicable to today's LLM/agent era.
More from Safety
- Utah becomes first US state to let AI prescribe medication without direct doctor review — NathanpmYoung · 2026-10-08
- Polymarket prices just 13% odds of a US AI safety bill by end of 2026 — Polymarket · 2026-10-08
- National Compute gifts $100M in compute credits to White House Genesis Mission — typewriters · 2026-10-08
- TheZvi: Automating Alignment Research Is Close to the Worst Possible Plan — TheZvi · 2026-10-08
- ChatGPT for Teens poses unacceptable risk as parental suicide alerts fail, new research finds — Polymarket · 2026-10-08
- Security researcher: the best AI bugs live at the safety-security intersection — wunderwuzzi23 · 2026-10-08