240k-edit study: humans fail to humanize AI-generated scientific writing, COLM 2026
sethlazar · x · 2026-10-08
A COLM 2026 paper ran 240k edits on scientific abstracts to test whether AI-assisted academic writing can be fixed.
Key findings:
- Humans are surprisingly bad at humanizing AI-generated scientific writing, often failing to erase human-AI differences.
- LLM editors don't help either—and frequently make already-strong abstracts worse.
Set against rising scrutiny of AI-assisted scientific writing, the results suggest post-hoc human editing is not a reliable fix for AI writing traces.
More from Research
- Stanford's Anshul Kundaje Slams 'Universal Virtual Cell': Perturbation Prediction Is the Wrong Metric — anshulkundaje · 2026-10-08
- Kernaut: coding agents design Gaussian process kernels via program search — sirbayes · 2026-10-08
- Kernaut's discovered kernels stay interpretable: 16 scalar functions, inner-product form — sirbayes · 2026-10-08
- iOSWorld brings computer-use agent benchmarking to iOS at COLM 2026 — kohjingyu · 2026-10-08
- Epoch's InnovationEval: AI agents still far from producing real research innovations — Afinetheorem · 2026-10-08
- Dankrad: AI's inelegant math results reveal human bias from small context windows — CatAstro_Piyush · 2026-10-08