Fudan and CUHK MMLab publish first survey of Agentic Visual Generation, classifying systems L0-L4
jiqizhixin · x · 2026-10-01
Fudan University's Wan Team and CUHK MMLab have released the first survey of Agentic Visual Generation, addressing the shift from "prompt in, image out" toward systems that autonomously decide what to generate, pick tools mid-execution, revise from results, and retain experience.
- The survey builds a classification framework around the controller's decision scope, grading systems from L0 to L4 by the farthest decision the controller can influence
- It covers image generation, editing, video, slides, UIs, 3D, and world models under one set of questions: what can the system decide before generation, choose during execution, change after seeing results, and retain after the task ends
- Key finding: current systems vary widely in autonomy across these dimensions, and the field lacks a shared language for where a generation pipeline ends and an agent begins
More from Multimodal
- Gemini 4 Pro Turned a Fly Into a Luxury Watch in One Prompt, Under 7 Minutes — cgarciae88 · 2026-10-01
- A Bedtime Story Made Entirely in Adobe Firefly: Art, Animation, Music and Voice — LudovicCreator · 2026-10-01
- MV for 'the last generation that wrote code by hand', every frame drawn by code — Relevant-Student-468 · 2026-10-01
- Interactive AI film STILL #1 fixes its character's gait and adds a rain effect — Daniel_Farinax · 2026-10-01
- Marigold V2 hands-on: running batch depth/normals on a dedicated GPU — Admirable-Cod-Lara · 2026-10-01
- User shares that Midjourney's recent article header images have been consistently impressive — AndyMasley · 2026-10-01