Visualizing a modular-addition transformer: 2k vs 20k training steps reveal ring structure
Sauers_ · x · 2026-09-24
A striking visualization compares a modular-addition transformer at 2k vs 20k training steps: 113 possible output tokens arranged clockwise, each ring a number being added, color showing logprobs. By 20k steps the model develops clear ring structure — a vivid window into how transformers learn modular arithmetic (the Grokking phenomenon).
Related event: Visualization Shows How a Modular Addition Transformer Learns(2 posts)→
More from Research
- Semantic operators: LLM data processing at scale needs full-stack rethink — CShorten30 · 2026-09-24
- Jev processes 100k rows for $2.50 in under 60s, demo now public — CShorten30 · 2026-09-24
- CLM-8B hits SOTA 81.6% on DeepSWE with light finetuning, up to 9x faster inference — anshulkundaje · 2026-09-24
- SpeakerMem-R1 Tops EverMemBench at 62.33% with Dual-Track Multi-Party Dialogue Memory — zju · 2026-09-24
- CMU's WhatWorkedBench Measures How Well AI Research Agents Understand Their Experiments — CarnegieMellonU · 2026-09-24
- Tencent's RewardVerse uses rubric-guided optimization to fix video reward model drift — tencent · 2026-09-24