Interpretability Experiment Collides Two Scene Memories

graceisford · x · 2026-07-13

This post recounts a series of mechanistic interpretability experiments on a world generation model. During the generation process, the author attempted activation patching, injecting intermediate layer activations from a generation of Monet's Poppy Field into the same layer during the generation of New York's Manhattanhenge.

The result visualized an effect of "two memories fighting for the same space." The author emphasizes that mechanistic interpretability can be integrated with artistic generation, using controllable intermediate interventions to observe how a model's internal representations influence its output.

Original post →

More from Research

Research channel →