Tester: AI-text detector Pangram shows zero false positives, but adversarial rewriting evades it

alex_peys · x · 2026-09-04

alexpeys reports hands-on testing of the AI-text detector Pangram: despite trying hard to break it, he has never seen a false positive.

He did manage false negatives by adversarially rewriting text against the classifier. His take: social norms will evolve into contexts where AI writing aid is or isn't acceptable, and good detection will matter.

Original post →

More from Models

Models channel →