Most open-source AI text detectors can't hold a 0.5% false-positive rate

grumpyp2 · reddit · 2026-09-02

A team benchmarked every notable open-source AI-text detector under one protocol—thresholds matched to 0.5% FPR on 6,930 human docs, then recall measured on raw AI, humanizer-paraphrased, and frontier-model text (GPT-5.x, Claude, Gemini 3.x):

Top scorer tropa-mini (ROC-AUC 0.968) is the authors' own, released as open weights (Apache-2.0) with full data and methodology for reproduction.

Original post →

More from Models

Models channel →