LLM version turnover breaks AI-writing screening: detectors miss 1 in 3 rewrites of newest models

rohanpaul_ai · x · 2026-10-10

A Tokyo Metropolitan University study (arXiv:2610.11599) quantifies how LLM turnover undermines AI-assisted writing screening in journals. The authors paired 4,000 pre-ChatGPT PNAS abstracts with rewrites by 23 LLM versions from three vendors (June 2023–August 2026) and trained detectors under maintenance scenarios from constant retraining to never updating.

Key findings:

The authors conclude that benchmarking detectors against fixed LLM versions is not sustainable as models keep changing.

Related event: AI Text Detectors Collapse as LLMs Iterate, Study Finds(3 posts)→

Original post →

More from Safety

Safety channel →