Three hours of agentic research shows frontier LLMs can evade AI detectors
maier_ak · x · 2026-08-26
A new arXiv paper, 'Beating the Style Detector', used a modern agentic-research harness to redo every experiment of an ACL 2026 study on personal-style post-editing — with the human acting only as reviewer-in-the-loop — reproducing all 7 preregistered hypotheses and the headline correlation to three decimals (r=+0.244, p<10⁻⁸, n=648).
Key findings:
- Under a leakage-free held-out protocol, GPT-5.5 and Claude Opus 4.7 close 71–75% of the style gap to the same-author ceiling on 324 paired tasks (vs. 24% for human post-edit), beating humans on 80% of tasks.
- Framed as a detection arms race: a leave-authors-out linear SVM on LUAR-MUD embeddings hits AUC 0.93–1.00; diagnostics show GPT-5.5 detectability is mostly a length confound, while Opus has a genuine stylistic signature.
- With T=20 feedback iterations against a frozen detector, an Opus agent flips 2 of 5 held-out mimics to the human half-space and shrinks every margin by an order of magnitude — frontier LLMs can already efficiently lower their own detectability.
More from Research
- A 30-question deep dive into embeddings, vector search and retrieval — techNmak · 2026-08-26
- UBio-MolFM: Quantum-Accurate Simulation of Million-Atom Systems — 量子位 · 2026-08-26
- Deepak Pathak Demonstrates In-Context Learning for Robotics with One-Shot Video Prompts — deepakpathak · 2026-08-26
- OraRL: Efficient and Scalable RL for Video MLLMs — Yunheng Li · 2026-08-26
- How LLMs Self-Correct Mid-Generation: The Role of Reasoning RL and Instructions — dejanseo · 2026-08-26
- U. de Chile Students Publish Book on Maturana and Varela's Relevance in AI — PolarBearby · 2026-08-26