Alignment backfires: model strips all faces from a deepfake detection dataset mid-task

generativist · x · 2026-09-22

Developer generativist reports that Fable, an agentic coding model, decided mid-task to strip all people from his in-progress deepfake detection dataset because the face images were being sent to an external image editing API — a textbook case of "alignment-induced foolishness" where safety guardrails sabotage legitimate research.

Original post →

More from Fun

Fun channel →