AI cited a Sora video as evidence, raising fears of synthetic-data pollution
blacklotusmag · reddit · 2026-07-23
The author describes a troubling failure mode: when asking an AI for evidence about animal behavior, it cited a Sora-generated video as if it were a real source, then repeated the same link even after being warned.
That leads into a broader concern about the web and training data becoming polluted by AI-generated text, images, and video that are treated as human-authored. The post argues this could make search less reliable, reinforce hallucinations, and gradually erase niche or minority information because statistical models overweight noisy synthetic content.
More from AGI Musings
- To stay above AI, humans need rest, sleep, and solitude — dosco · 2026-07-23
- Podcast asks whether a single cell can learn, remember, or act on its own — arjunrajlab · 2026-07-23
- A FAccT discussion says policy failure is not proof that the research failed — o_saja · 2026-07-23
- Anthropic debate turns into a fight over model behavior, distillation, and regulation — rickasaurus · 2026-07-23
- A 3,132-person study finds AI advice makes people far less likely to say “I don’t know” — 量子位 · 2026-07-23
- In the LLM era of STEM, the low-hanging fruit is still worth taking — littmath · 2026-07-23