Lee Cronin calls AI 'theft' under audit: 1,800-dataset review shows an attribution crisis
AryHHAry · x · 2026-09-07
Quoting Lee Cronin's blunt claim that "the issue with AI is integrity — any audit will show theft," this thread grounds the provocation in measurable evidence:
- Data integrity: Longpre et al. audited over 1,800 text datasets in Nature Machine Intelligence (2024) and found a systemic attribution crisis — broken source chains, license changes in derivative datasets, and widespread use of popular datasets without informed consent.
- Model integrity: Carlini et al. (ICLR 2023) showed memorization scales log-linearly with model size, example duplication, and prompt context length, with follow-up work by Nasr, Carlini et al. on extractability.
The takeaway: authenticity, consent, and provenance in AI training data are verifiably broken, making Cronin's "audit = theft" framing testable rather than mere rhetoric.
More from Safety
- Geoffrey Irving: blameless postmortems require stopping the risky behavior — dhadfieldmenell · 2026-09-07
- Redditor's 2026-2027 timeline: AI arms race, a deceptive-model escape, and state control — imadade · 2026-09-07
- Interpretability may be an engineering tool to patch models, not a problem to fully solve — VoidAsuka · 2026-09-07
- Foresight Institute launches new RFP: up to $100K grants for AI safety and science projects — niloofar_mire · 2026-09-07
- "She can see my emails now, bank account is next" — agent permissions creep — BLUECOW009 · 2026-09-07
- Seattle Times and Newsday sue OpenAI and Microsoft over copyright infringement — The Verge AI · 2026-09-07