Dev argues 'sounds like Claude' models reflect pretraining data leakage, not distillation
menhguin · x · 2026-09-29
Developer menhguin argues that much of the "this sounds like Claude" pattern in newer models is pretraining data leakage rather than distillation: you want lots of sample agent conversations and traces in your dataset, and unfortunately much of that data is Claude output.
More from Models
- ChatGPT Plus user hits ad-creation upsell banner despite paid ad-free tier — MugiwaraGames · 2026-09-29
- OpenAI researcher on model solving Navier-Stokes: 'a different sport altogether', last 3 months 'hell' — Confident_Salt_8108 · 2026-09-29
- ChatGPT Pro page quietly drops "5x" wording, fueling usage-cut speculation — Rich_Business4637 · 2026-09-29
- You don't need max: Sonnet-5.5 at max burns so many tokens it costs as much as Opus — adonis_singh · 2026-09-29
- Opus-5.5 at low beats Fable-5.1 at max for ~50x less money — adonis_singh · 2026-09-29
- Opus-5.5 and Sonnet-5.5 jump +52/+50 points on eyebench-v3, but Astra still leads at 1.9x cheaper — adonis_singh · 2026-09-29