Marionette predicts world states and renders geometry via diffusion
Zian Meng · hf · 2026-08-17
Marionette predicts explicit 3D articulated world states for games, uses a fixed renderer for geometry, and synthesizes video via diffusion. This enables direct state-level control and long-horizon consistency repair.
More from Multimodal
- 60-Year-Old Chinese Grandpa Creates 20-Minute Cyberpunk Anime with AI, Praised as 'Game CG Quality' — SimplyAnnisa · 2026-08-17
- New Benchmark Reveals AI-Generated Video Detectors Fail on Real-World Crisis Events — huggingface · 2026-08-17
- MiniMax H3 generates otter video locally in 3 minutes — emollick · 2026-08-17
- Yinchao V4 Released: Architecture Rebuild Solves Chinese Singing Challenges — 新智元 · 2026-08-17
- Seedance 2.5 Tops Multi-Image-to-Video Benchmark with Elo 1400 — rohanpaul_ai · 2026-08-17
- Takeaway: Image-to-Video Outperforms Text-to-Video in Local Tests — Far-Solid3188 · 2026-08-17