Tester: AI-text detector Pangram shows zero false positives, but adversarial rewriting evades it
alex_peys · x · 2026-09-04
alexpeys reports hands-on testing of the AI-text detector Pangram: despite trying hard to break it, he has never seen a false positive.
He did manage false negatives by adversarially rewriting text against the classifier. His take: social norms will evolve into contexts where AI writing aid is or isn't acceptable, and good detection will matter.
More from Models
- Gary Marcus asks: is Astra a pure LLM or an undisclosed neurosymbolic hybrid? — GaryMarcus · 2026-09-04
- GPT-6 Astra hits 62.7% on ARC-AGI-3, 99.9% with new adapter harness, sets ARC-AGI-2 SOTA at 95% — BlackHC · 2026-09-04
- Users burned rate limits expecting GPT-6 Astra reset, but launch slipped to Friday — BLUECOW009 · 2026-09-04
- Gary Marcus blocks critic amid Astra architecture fight: 'nothing has been published' on how it works — GaryMarcus · 2026-09-04
- OpenAI ships GPT-6 Astra, first model to trigger internal safety measures, billed as AGI — ShakeelHashim · 2026-09-04
- OpenAI unveils GPT-6 Astra: an agent that can do anything on your computer — AndrewSchmidtFC · 2026-09-04