Why “it’s just distilled from GPT/Claude” is often weak evidence, not proof

UsedMorning9886 · reddit · 2026-07-23

This Reddit post argues that accusations of model “distillation” are often overstated.

The author distinguishes true token-level distillation — which requires logits — from the much more common practice of training on text completions from larger models. They argue that public APIs only expose filtered text, not logits, so many smaller teams are really doing synthetic-data bootstrapping rather than literal distillation. The post also notes that similar synthetic-data practices are used across the industry, including by closed labs training on their own earlier model outputs, and suggests that “it’s just distilled from GPT/Claude” is often weak evidence rather than proof.

Related event: AI Distillation and IP Boundary Debate Sparks Anthropic Backlash(13 posts)→

Original post →

More from Research

Research channel →