Dev predicts a universal DSL for video data will enable highly controllable world-simulation diffusion models
zeeshanp_ · x · 2026-09-30
Zeeshan Pishevar argues LLMs are now good enough to create videos via code, so video descriptions can be expressed in code too — building on prior practice of JSON-formatted synthetic captions for generative visual models. He predicts a universal DSL for video data will arrive soon, enabling diffusion models trained for extreme controllability in world simulations.
More from Multimodal
- Dolphin AI launches agentic video studio with multi-shot character consistency, 45k test videos — thetripathi58 · 2026-09-30
- Dioramas open-sources a free 3D website framework with AI-generated assets and 20 example sites — Scobleizer · 2026-09-30
- Niantic Spatial demos a short film built inside a real-world Gaussian splat in hours — Scobleizer · 2026-09-30
- AI-reanimated Greta Garbo stars in SKF ball-bearing ad, panned as bland — nordicinst · 2026-09-30
- MiniMax H3 at max settings takes 26 min and 62GB VRAM per 15s clip on RTX Pro 6000 — Realistic-Fennel-190 · 2026-09-30
- VideoLoop rewrites bounded working memory, hits 88.3% on VideoMME long video — Jinfa Huang · 2026-09-30