LLMs Decode Early MDLM Gibberish: Predict The Verge as Training Data Source

dejanseo · x · 2026-08-02

A developer observed that the output from an early checkpoint of a masked diffusion language model (MDLM) exhibited a distinct "hallucinated tech journalism" style. Upon feeding the text to Gemini, GPT, and Claude, all three models consistently identified The Verge as the original training data source.

The LLMs successfully recognized topic overlaps—such as consumer tech, VR, gaming, and space—as well as specific stylistic artifacts characteristic of the publication. This highlights the potential of LLMs in tracing and identifying stylistic data lineages.

Original post →

More from Models

Models channel →