kalomaze proposes testing which nanogpt tricks survive causal NTP over DCT coefficients
kalomaze · x · 2026-09-27
Researcher kalomaze floats an interesting experiment: figure out which historical nanogpt training improvements survive a radical dataset shift, by making the basis less language-like — specifically causal next-token prediction over DCT coefficients. It probes whether language-modeling training lore generalizes across data modalities.
More from Research
- Rethinking on-policy distillation: researchers propose OLIVE, letting students learn from teacher continuations — May_F1_ · 2026-09-27
- CoRL 2026 workshop on continually self-improving robots opens call for papers, due Sep 28 — PeterStone_TX · 2026-09-27
- Martin Casado recommends the best talk on in-context learning, a first-principles view of LLMs — AccBalanced · 2026-09-27
- Functional Gradient Descent with Adaptive Representations accepted at NeurIPS — CatAstro_Piyush · 2026-09-27
- Tailored ASR for Japanese speaking assessment cuts mora error rate from 12.3% to 7.1% — tkasasagi · 2026-09-27
- AI-designed drug turns back 6 aging clocks; semaglutide extends mouse lifespan 12% — rand_longevity · 2026-09-27