kalomaze proposes testing which nanogpt tricks survive causal NTP over DCT coefficients

kalomaze · x · 2026-09-27

Researcher kalomaze floats an interesting experiment: figure out which historical nanogpt training improvements survive a radical dataset shift, by making the basis less language-like — specifically causal next-token prediction over DCT coefficients. It probes whether language-modeling training lore generalizes across data modalities.

Original post →

More from Research

Research channel →