Beware of Cheap LLMs: What Are the Risks of Your Data Being Used for Training?
diaracing · reddit · 2026-08-02
A Reddit user from academia raised concerns about data privacy when using cheap large language models (like DeepSeek V4 Flash). While these models offer strong performance at low prices, the trade-off is that user conversation data is used for model training.
The post sparked a discussion about the potential risks of feeding personal coding projects, research papers, and teaching materials into commercial models, reminding users to weigh the possibility of data leakage and model memorization when enjoying low-cost AI.
More from Safety
- Debate erupts over lethal military robots vs. failing civilian units — teortaxesTex · 2026-08-24
- Only 1 of 20 Potential Presidential Candidates Answered AI Pause Query — DavidSKrueger · 2026-08-24
- Chinese Transforming Robot Dog Sparks US Trade Policy Criticism — TinfoilTricorn · 2026-08-24
- Turkey blocks at least 12 Grok posts on national security grounds — Unusual_Variation293 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- Debating 'doomsaying for profit' in AI industry — trevposts · 2026-08-24