Deep Dream and SAE Feature Maximization in Modern LLMs

torchcompiled · x · 2026-07-20

A reply to @matthen2's discussion on the Deep Dream effect in modern large models. The original post demonstrated optimizing an image to maximize the probability of a target caption. It found that Gemma 12B, which lacks a vision encoder, reads pixels just like processing token embeddings, printing recognizable objects on the canvas. Conversely, E4B, which features a vision encoder, tends to lean towards texture drift.

The responder noted that maximizing optimization using individual features identified by Sparse Autoencoders (SAE) might yield even more interesting results.

Related event: Reviving Deep Dream with modern LLMs(2 posts)→

Original post →

More from Research

Research channel →