"Trust No AI": Security Researcher Revives 2020 "Assume Bias" Playbook for AI Safety

wunderwuzzi23 · x · 2026-10-02

Security researcher wunderwuzzi (Embrace The Red) is echoing arekfurt's call for an "assume misalignment" concept in AI safety, analogous to cybersecurity's "assume breach" principle. He proposed the same idea six years ago as "Assume Bias" in his ML attack series, citing failures like Amazon's biased recruiting AI, Microsoft's Tay, and IBM's cancer-treatment recommender.

The core argument: just as network security assumes systems are already compromised and plans detection/response accordingly, AI architects should assume models will be biased and produce wrong predictions — building in mitigation, testing, detection, and response (including deprecation) strategies from the start. He now distills the mindset into a simpler slogan: "Trust No AI."

Original post →

More from Safety

Safety channel →