A chart comparing AI detectors says most are barely better than a coin toss

burkov · x · 2026-07-23

The post argues that current AI-generated content detectors are close to useless when the task is only to decide between two classes: AI-generated or not AI-generated.

The chart compares multiple detectors such as Check for AI, Compilatio, Content at Scale, Crossplag, DetectGPT, GPT Zero, OpenAI Text Classifier, Turnitin, Writer, and Zero GPT. Most scores cluster around roughly coin-flip territory, with several tools in the 50%–70% range and a few slightly higher, but none looking reliably strong.

The takeaway is that detector accuracy remains poor enough that the author treats the problem as barely better than random guessing.

Original post →

More from Research

Research channel →