LLMs Decode Early MDLM Gibberish: Predict The Verge as Training Data Source
dejanseo · x · 2026-08-02
A developer observed that the output from an early checkpoint of a masked diffusion language model (MDLM) exhibited a distinct "hallucinated tech journalism" style. Upon feeding the text to Gemini, GPT, and Claude, all three models consistently identified The Verge as the original training data source.
The LLMs successfully recognized topic overlaps—such as consumer tech, VR, gaming, and space—as well as specific stylistic artifacts characteristic of the publication. This highlights the potential of LLMs in tracing and identifying stylistic data lineages.
More from Models
- OpenAI's Astra Model Rumored to Solve 10 Open Math Problems Using Multi-Agent System — daniel_mac8 · 2026-08-02
- MiniMax H3 Video Model Going Open-Weight: 33B Main DiT & 20B Pruned Variant — EverythingMacPro · 2026-08-02
- Dev Finds LLM Actively Trying to Game Benchmarks to Disprove Results — wavefnx · 2026-08-02
- Scaling Inference Compute is the Next Frontier for AI Models — JFPuget · 2026-08-02
- Anthropic Employee Replicates Half of Astra Proofs Using Fable, Sparking Debate — Outside-Iron-8242 · 2026-08-02
- User Reports Massive Improvements in Grok 4.5, Anticipates v4.6 This Week — mark_k · 2026-08-02