kalomaze on long-sequence training: conditional uncertainty and information asymmetry explain gradient absorption

kalomaze · x · 2026-10-07

In a discussion on how gradients are absorbed in long-context training, kalomaze argues the root cause is conditional uncertainty and information asymmetry. NLL of future tokens improves on average as sequence length grows because the model gains more conditional information, and ICL absorbs conditioning gradients in a way skewed toward late-token improvement. This responds to questions about attention-head distribution and whether early high-frequency patterns inhibit hierarchical modeling—needing validation at greater scale.

Original post →

More from Research

Research channel →