Token superposition pretraining idea: prior experiments were 'pretty much a disaster'
giffmana · x · 2026-09-19
- Aleksa Gordic proposes TST: during early pretraining, group every s neighboring tokens and average their embeddings into one latent token that predicts the next non-overlapping bag of s tokens (inspired partly by resolution in vision).
- Kevin (giffmana) replies he ran related experiments inspired by visual resolution — most turned out to be "pretty much a disaster".
- A rare honest negative-result datapoint for multi-token training ideas.
Related event: Token Superposition Training Could Cut MoE Pretraining Compute by 2.5x(2 posts)→
More from Research
- Google open-sources Fuse, a multi-agent framework for verifiable social reasoning in LLMs — google · 2026-09-19
- Study finds scientific beauty and impact are surprisingly correlated — jacobkimmel · 2026-09-19
- Trolley problem: Jev-style API turns jina-reranker-v3.5 into a ruthless decision engine — gaganghotra_ · 2026-09-19
- Probe guidance steers continuous diffusion LMs with just 1-3% extra inference compute — itsbautistam · 2026-09-19
- LLM Analysis of Wikipedia Battle Pages Crowns Napoleon the GOAT With 16.7 Career WAR — ctjlewis · 2026-09-19
- MIT paper tracks the birth, life and death of 108k open-weight AI models on Hugging Face — asusarla · 2026-09-19