One prompt, ten panels: Reddit's recursive zoom image-generation experiment
NVDA808 · reddit · 2026-09-16
A Reddit user shared a single-prompt image generation experiment: the model produces one large image containing 10 sequential panels, where each panel must be an enlarged interpretation of a real region of the previous panel.
Key prompt design:
- Panel 1: generate a richly detailed original scene full of explorable visual information, with no pre-planned path.
- Panels 2–9: after each panel, the model inspects its own output, marks the most interesting genuine feature with a crop box, and treats only that region as the structural parent of the next frame — a strictly nested recursive zoom.
- No lazy zooms: each step must reveal a new level of organization, avoid repetitive microscopic textures, and never invent elements with no visual precursor.
The whole loop of generate-observe-select-zoom is executed autonomously by the model with zero user intervention, framed as a 'visual contact-sheet exploration.'
More from Multimodal
- Creative short film made with Seedance 2.5 shows cinematic crowd-freeze scene — SimplyAnnisa · 2026-09-17
- Generative AI short film submitted to Lumara Film Festival — Kyrannio · 2026-09-17
- Google releases Gemma 3n: 2GB RAM multimodal model, first sub-10B to top 1300 on LMArena — joemeno · 2026-09-17
- GPT-6 Astra tested on complex traditional architecture, large-scale layout holds up well — Due-Emu7804 · 2026-09-17
- Testing GPT-6 Astra on Traditional Architecture: Symmetry Holds Up — Due-Emu7804 · 2026-09-17
- NetEase Youdao open-sources Confucius R2T2, a 2B speech model with 200ms real-time transcription — dr_cintas · 2026-09-17