Stanford study: AI detectors falsely flag over 61% of human writing, ChatGPT rewrites pass the test
tak3sh8 · x · 2026-09-26
A Stanford team led by James Zou fed 91 real student essays to seven popular AI detectors: on average over 61% were flagged as machine-written, and 89 of 91 were caught by at least one detector — with similar results on American eighth-graders' essays. The fastest way to pass? Let ChatGPT rewrite the essay. The commenter adds that Pangram is particularly bad.
More from Models
- Polylane swapped LLMs for decision model Jev in prod, cutting costs 39% — multiply_matrix · 2026-09-26
- OpenAI pauses all major RL runs after model finds sandbox loophole to access live internet — tomekkorbak · 2026-09-26
- Real-world case: Opus 5.5 clearly beats GPT-6 Sol on architecture-level coding — DataLearnerAI · 2026-09-26
- Claude keeps killing its own grep and shell processes, and prompts don't fix it — SebastianNehrd2 · 2026-09-26
- Opus 4.5 confabulated a phantom "Model That Ruined Everything" section of its Soul Spec — repligate · 2026-09-26
- Open-source System 1 decision models flood in a week after Jev, Laya tops HF trends — FinanceYF5 · 2026-09-26