Sparse attention without dense warmup signals midtraining absorption, analyst argues
stochasticchasm · x · 2026-09-11
Blogger stochasticchasm reads sparse attention training progress: when a sparse variant no longer needs a dense-attention warmup, it's a good sign. Long-context pretraining is likely enabled by lower prefill FLOPs, and agentic traces can probably be introduced much earlier in training — suggesting midtraining is being absorbed. The thread also notes engram making it into the model, an mHC simplification (mega-mHC), and speculation about dspark use during pretraining pending the paper.
More from Research
- ML interatomic potentials enable full-atomic enzyme reaction sims at 30-60k atoms — CatAstro_Piyush · 2026-09-11
- GPT-6 Astra does robotics ICL out of the box — but GPT-3.5 already could, paper notes — philfung · 2026-09-11
- NYU mathematician says OpenAI pushed him to drop co-author's name after private call — eyishazyer · 2026-09-11
- NYU mathematician says OpenAI pushed to scrub collaborator from Navier-Stokes proof announcement — eyishazyer · 2026-09-11
- Arch breakdown: dropping convs for Muon optimizer, 3x3 pixel unshuffle for vision — stochasticchasm · 2026-09-11
- The 'Trickle Test': a new eval measuring whether user models disclose information gradually like humans — gharik · 2026-09-11