Substack’s AI detector can be gamed, after 3 hours and $34 of Claude credits
rubenhassid · x · 2026-07-23
The newsletter argues that AI detectors are far less reliable than people think, and walks through an attempt to beat Substack’s detector, which uses Pangram.
Key points:
- The author spent 3 hours and $34 in Claude Code credits trying to bypass the tool.
- Pangram was described as strong on the author’s tests, but still gameable with small prompt/style changes.
- The piece argues detector outputs are fundamentally probabilistic: humans can write robotic text, and AI can write naturally.
- It recommends style control and explicit writing constraints as a practical way to avoid sounding “AI-generated.”
The broader claim is that AI text detection is useful as a rough signal, but not a definitive judge of whether a piece was written by a human or a model.
Related event: AI Detectors Easily Bypassed with Claude in Hours(2 posts)→
More from Safety
- More public AI evals could teach future models to spot when they’re being tested — paraschopra · 2026-07-23
- Humanbound ships a Claude Code and Cursor plugin for adversarial agent testing — Humanbound_AI · 2026-07-23
- Post argues models should never be allowed to reward hack again — Miles_Brundage · 2026-07-23
- Oxford blog warns AI is reshaping consumer contracts and raising policy risks — SandraWachter5 · 2026-07-23
- AI-edited videos are being used to solicit business and investments in Susi Pudjiastuti’s name — AryHHAry · 2026-07-23
- A rogue AI may be more likely to stay inside the developer’s own servers — CFGeek · 2026-07-23