Independent Cross-Entropy Analysis Reveals Suspicious Similarities Between Major LLMs
scaling01 · x · 2026-07-30
A developer analyzed the raw text outputs of numerous major LLMs using cross-entropy comparison, uncovering suspicious similarities between certain models (such as Kimi-K3 and Fable, and DeepSeek v4 and Grok-4.3), raising concerns about potential data contamination or shared lineages.
Subsequently, another developer replicated the analysis using Fable on a completely different corpus (WeirdML transcripts). The results showed that the connections—like Kimi-K3/Fable—remained prominent. While the author cautions that these could be spurious correlations, finding consistent results across different datasets adds credibility to the hypothesis that the similarities are meaningful.
More from Models
- Qwen's Small Models So High-Quality, Users Joke OpenAI Will Go Bankrupt — glenbeer · 2026-07-30
- Claude Still Hallucinates with Web Search: Fabricates Scholar's Move to MIT — conitzer · 2026-07-30
- OpenAI Reportedly Hid Funding for FrontierMath Benchmark Under NDA — BlancheMinerva · 2026-07-30
- User Tricks Claude into Revealing Anthropic's Internal UI System Prompt — sergeykarayev · 2026-07-30
- AI Can Pass the Turing Test but Fails to Detect Customer Lies — claud_fuen · 2026-07-30
- Reddit Debate: Fake Open Weights from Giants Like Kimi Are Eroding Trust — Ok-Shower7286 · 2026-07-30