Interpretability Experiment Collides Two Scene Memories
graceisford · x · 2026-07-13
This post recounts a series of mechanistic interpretability experiments on a world generation model. During the generation process, the author attempted activation patching, injecting intermediate layer activations from a generation of Monet's Poppy Field into the same layer during the generation of New York's Manhattanhenge.
The result visualized an effect of "two memories fighting for the same space." The author emphasizes that mechanistic interpretability can be integrated with artistic generation, using controllable intermediate interventions to observe how a model's internal representations influence its output.
More from Research
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21
- AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop — 404 Media · 2026-07-21
- Shared agent workspaces fail in a fixed order, from stale reads to zombie writes — mrvladp · 2026-07-21
- Practical rolling-shutter pose estimation uses affine correspondences — ducha_aiki · 2026-07-21