Beware of Cheap LLMs: What Are the Risks of Your Data Being Used for Training?

diaracing · reddit · 2026-08-02

A Reddit user from academia raised concerns about data privacy when using cheap large language models (like DeepSeek V4 Flash). While these models offer strong performance at low prices, the trade-off is that user conversation data is used for model training.

The post sparked a discussion about the potential risks of feeding personal coding projects, research papers, and teaching materials into commercial models, reminding users to weigh the possibility of data leakage and model memorization when enjoying low-cost AI.

Original post →

More from Safety

Safety channel →