Testing Text Detectors With Pre-Internet Books
cephaloform · x · 2026-07-17
- This retweet discusses how, when evaluating text detectors like Pangram, you can't just test them on public texts from before the ChatGPT era, or the results will be contaminated by the fact that "the model has seen the training data."
- To verify it isn't just memorizing old text, someone found old books that don't exist on the internet, scanned them, and ran tests totaling nearly 3 million words.
- In the results, the detector still found a few traces of AI: 3 sections suspected of AI assistance, and 1 section suspected of AI generation.
- The key takeaway isn't the conclusion itself, but the rigorous anti-"data leakage" evaluation approach it provides.
More from Research
- MaP-WAM tackles non-Markovian robot manipulation with memory-grounded planning — Sizhe Zhao · 2026-09-11
- Negative Self-Distillation improves LLM reasoning by avoiding flawed reasoning paths — Rongcan Pei · 2026-09-11
- DeepMind-led paper makes design docs the source of truth, code disposable — SMART regenerates in 1.5-3h for ~$100 — Roger_M_Taylor · 2026-09-11
- GameWorld wins Best Paper Runner-Up at ECCV 2026 Multimodal Digital Agents Workshop — MikeShou1 · 2026-09-11
- Yann LeCun live at ECCV on World Models — Weak_Assistance_5261 · 2026-09-11
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11