Is there any LLM whose training data is fully auditable and un-stolen?
LuCiAnO241 · reddit · 2026-09-25
The poster asks whether any LLM can claim its training data is entirely un-stolen. Their search led to fully open-data models like OLMo, whose datasets can be inspected and audited — though the author admits they wouldn't know how to audit them. They wonder if any model actually prides itself on clean data, joking that aside from a 1930s public-domain-era model, the answer seems to be no.
More from Models
- After RL training, calling LLMs 'language predictors' is no longer accurate, researcher argues — morqon · 2026-09-25
- Anthropic and OpenAI swap playbooks: generous usage vs. user-hostile limits — OwariDa · 2026-09-25
- Dev claims Codex is 10x less token-efficient than Claude Code: $20 buys one day vs one week — DimitrisPapail · 2026-09-25
- Passed-around take: you're better off treating LLMs as brute-force tools — burny_tech · 2026-09-25
- Ego promo video shows brutal model comparison as Claude Opus 5.5 impresses — vista8 · 2026-09-25
- Opus 5.5 takes #1 on CADArena at 0.750, one-shots 3D animation — hudzah · 2026-09-25