Is there any LLM whose training data is fully auditable and un-stolen?

LuCiAnO241 · reddit · 2026-09-25

The poster asks whether any LLM can claim its training data is entirely un-stolen. Their search led to fully open-data models like OLMo, whose datasets can be inspected and audited — though the author admits they wouldn't know how to audit them. They wonder if any model actually prides itself on clean data, joking that aside from a 1930s public-domain-era model, the answer seems to be no.

Original post →

More from Models

Models channel →