A chart comparing AI detectors says most are barely better than a coin toss
burkov · x · 2026-07-23
The post argues that current AI-generated content detectors are close to useless when the task is only to decide between two classes: AI-generated or not AI-generated.
The chart compares multiple detectors such as Check for AI, Compilatio, Content at Scale, Crossplag, DetectGPT, GPT Zero, OpenAI Text Classifier, Turnitin, Writer, and Zero GPT. Most scores cluster around roughly coin-flip territory, with several tools in the 50%–70% range and a few slightly higher, but none looking reliably strong.
The takeaway is that detector accuracy remains poor enough that the author treats the problem as barely better than random guessing.
More from Research
- RECAP lets probes verify activation explanations that reconstruction scores can fake — Hiskias Dingeto · 2026-07-23
- USC benchmark shows GPT-5.5 scores 10.6% on active visual observation tasks — UniversityofSouthernCalifornia · 2026-07-23
- Paper proposes Linearizing Softmax Attention into Gated DeltaNet — bronzeagepapi · 2026-07-23
- Overcoming RGB Sim-to-Real Gap: Gaussian Splatting Enables Zero-Shot Robot Transfer — ZeYanjie · 2026-07-23
- Long-context models still copy irrelevant text, and a new reward lifts accuracy by up to 4.6 points — rohanpaul_ai · 2026-07-23
- Real Task Cost Across GPT, Claude, Gemini, Kimi: 10.6x Spread Despite Only 2x Price Difference, Hidden Reasoning Tokens Blamed — pixelo2323 · 2026-07-23