Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific Writing
Kazuki Nakajima, Takayuki Mizuno
cs.CL, cs.AI, cs.CY, cs.DL, cs.SI
2026-10-08
AI-writing screens collapse at LLM generation boundaries: detector catch rates fall from 99% to 3.8%, and a commercial detector misses 80% of the newest model's rewrites.
Journals and conferences now screen submissions with automated AI-text detectors, and the stakes are concrete. In 2026, a leading computer science conference ran a commercial detector over every submission to one of its tracks and rejected 178 papers without review, 18.4% of that track. NeurIPS partnered with the detector Pangram to screen its 2026 position paper track. Every such decision assumes that benchmark accuracy carries over to live screening. Benchmarks, though, are built from the LLM versions available when a detector was evaluated, while the versions authors actually use keep turning over. Nobody had quantified what that turnover does to screening under the maintenance conditions an institution can realistically afford.
The design is clean. Take 4,000 PNAS abstracts published between 2009 and 2018, all predating ChatGPT and sampled evenly across four disciplinary domains, and have 23 LLM versions rewrite each one with a two-stage prompt: compress the abstract into key points, then expand it back into a single abstract paragraph. The versions span June 2023 to August 2026, covering OpenAI (14, GPT-4 through GPT-5.6 Terra), Meta (4, Llama 3.1 through Muse-Glimmer), and Alibaba (5, Qwen2.5 through Qwen3.8), leaving 91,953 retained rewrites.
Detectors are fine-tuned RoBERTa-base classifiers (a 125-million-parameter pretrained language model), with thresholds calibrated on 30,925 human abstracts no detector sees in training, holding the false positive rate at 1%. The key variable is maintenance, in four scenarios: train on the target version itself (current version, the ideal); retrain the moment a new version appears (past and current); train only on the vendor's earlier versions (past only, the common situation where a new model's text arrives before any training data exists); and train once on the earliest version and never update (frozen). Two screening architectures sit on top: a union of 23 version-specific detectors that flags whenever any one fires, and a single pooled detector trained on all versions. The commercial product Pangram is tested through its API at production settings.
| Approach | Metric | Result |
| Current version (median over 23 versions) | TPR at 1% FPR | 99.9% |
| Past and current versions | TPR at 1% FPR | 99.3% |
| Past versions only | TPR at 1% FPR | median 95.4%; 26.6% for GPT-5, 3.8% for Muse-Glimmer |
| Frozen at GPT-4, tested on GPT-5 onward | TPR | at most 0.6% |
| GPT-4o detector on Llama 3.3 | TPR | 99.5% (transfers across vendors in one generation) |
| GPT-4 detector on GPT-5.4 | TPR | 0.2% (fails across generations of the same vendor) |
| Union of 23 detectors | human abstracts flagged | 11.9%, one in eight |
| Pooled detector | Muse-Glimmer missed | 33.8%, one in three (median miss rate 2.8%) |
| Pangram | FPR / Muse-Glimmer missed | 0.02% (1 of 5,000) / 79.8% |
Four structural findings:
For editors and research-integrity officers the conclusion is concrete: a detector's benchmark accuracy is provisional and must be re-verified at every LLM release, with earlier versions retested too. The JSD signal is practical because it needs only a sample of rewrites from the new model, no parameters or architecture details, which for major closed models are unavailable anyway.
For detector builders, the lesson is that a high score on a fixed model pool says little about deployed reliability. The failure mode is generation boundaries, which is more actionable than the vague sense that newer models are harder to detect.
The authors list the main ones: one journal (PNAS), one language, one genre (abstracts); LLM use simulated as a fixed two-stage rewrite, while real assistance ranges from grammar fixes to full rewrites over many exchanges; one detector family (fine-tuned RoBERTa-base) plus one commercial product at a single operating point; three vendors, with Anthropic excluded over its usage terms; and the JSD-detection link is correlation, not causation. Two more on close reading: four public detectors without in-domain training, Binoculars and FastDetectGPT among them, missed over two-thirds of every version's rewrites, so the in-domain trained setting here already favors detection and deployed collapse could be worse; and Pangram was evaluated at version 3.3.2, which the vendor retired on September 30, 2026, so newer behavior is unknown.