Commercial detector re-runs NeurIPS AI-text check: 7.3% of 2025 papers flagged vs Pangram's 1%
AltruisticCouple3491 · reddit · 2026-10-08
- A developer of commercial AI-text detector ParaTrace re-ran the AI check NeurIPS did with Pangram 3.3.2 on 2022 (pre-ChatGPT) and 2025 position papers.
- 2022 control: neither detector flags any paper; ParaTrace scored only 0.4% of 4,512 passages as AI, with a 95% upper bound of 3.5% paper-level false positives.
- 2025 disagreement: at ≥90% both flag 2 papers, but at ≥50% ParaTrace flags 9/123 (7.3%) vs Pangram's 2/204 (1.0%).
- Four possible explanations: Pangram missing AI-assisted writing; ParaTrace by design flagging AI-polished human text (42.8% of AI-paraphrased human text is flagged, while NeurIPS allows AI copy-editing); false positives on 2025 human writing; chunk-size differences.
- On Pangram's public test sets, ParaTrace matches on clean AI text (100% on fully AI essays), trails on short texts (73% vs 99.7% under 50 words) and collapses on humanizer output (21% vs 99%), with higher default false positives (1.86% vs 0%).
- Takeaway: no ground truth exists for which 2025 papers used AI, detectors diverge sharply on gray areas like AI polishing, and acting on such flags is risky.
More from Safety
- Abusers hop across inference providers, so providers must coordinate evictions — natolambert · 2026-10-08
- Three fired OpenAI safety researchers demand board preserve chain-of-thought monitoring — mark_k · 2026-10-08
- Guardian probes AI workplace surveillance in Europe as Multiverse staff call monitoring 'unnerving' — nordicinst · 2026-10-08
- Reward hacking may point to a deeper misalignment problem in models, researcher warns — gleech · 2026-10-08
- Revisiting 'Specification Gaming Examples in AI', the 2019 resource born in an evidence vacuum — gleech · 2026-10-08
- Researchers manually tag 610 reward hacking cases; over 100 show no obvious environment flaw — gleech · 2026-10-08