Astra shifts video generation from pixel diffusion to building Blender 3D scenes for rendering

IndraVahan · x · 2026-09-07

IndraVahan announced a major leap with Astra: instead of diffusion models building video pixel-by-pixel, the system constructs Blender 3D scenes that are then rendered as video. The argument: pixel/frame generation always leaves traces of AI generation, while 3D models are more solid and give far greater control over scenes and every asset within them. This reframes "multimodality" as actually building assets to render frames, with implications for video generation, world models, gaming, and animation. The author also believes platforms like Meshy will play a big role in this 3D-first approach.

Related event: Astra ditches diffusion, generates video via 3D scene construction(2 posts)→

Original post →

More from Multimodal

Multimodal channel →