kalomaze on long-sequence training: conditional uncertainty and information asymmetry explain gradient absorption
kalomaze · x · 2026-10-07
In a discussion on how gradients are absorbed in long-context training, kalomaze argues the root cause is conditional uncertainty and information asymmetry. NLL of future tokens improves on average as sequence length grows because the model gains more conditional information, and ICL absorbs conditioning gradients in a way skewed toward late-token improvement. This responds to questions about attention-head distribution and whether early high-frequency patterns inhibit hierarchical modeling—needing validation at greater scale.
More from Research
- NeuroAI manifesto: applying AI scaling laws to BCIs, the bitter lesson for the brain — w1kke · 2026-10-07
- LeanLean benchmark: Opus 5.5 scores 64.3% compressing Lean proofs, GPT 6.1 Sol only 39.9% — ChrSzegedy · 2026-10-07
- PersistBench (NeurIPS Spotlight): 4D foundation models can see but not remember — weichiuma · 2026-10-07
- COLM 2026 poster: Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer — boknilev · 2026-10-07
- AI-Read Gold Electrodes Detect Molecular Chirality One Molecule at a Time — Brighter-Side-News · 2026-10-07
- BinkBench: A No-Cap Agent Benchmark for Video Quality and Compression, Seeking Testers — -MaskNinja- · 2026-10-07