Six years on, wunderwuzzi's ML attack series on model backdoors and Assume Breach still holds up

wunderwuzzi23 · x · 2026-10-08

Security researcher wunderwuzzi (Embrace The Red) revisits his 2020 Machine Learning Attack Series, written to bridge the gap between ML safety research and security practice. His core stance: adopt an Assume Breach mindset and ask "how would we even know it happened, and what do we do about it?"

The series uses a self-built image classifier (Husky AI) to cover ML threat modeling and hands-on attacks:

Notably, he observes these analyses remain highly applicable to today's LLM/agent era.

Original post →

More from Safety

Safety channel →