Unreal Engine pipeline yields 8.7K+ hours of action-conditioned video for world-model pretraining
udmrzn · x · 2026-09-06
- An arXiv paper presents an Unreal Engine-based pipeline generating 8.7K+ hours of synchronized multi-view video with frame-aligned actions and camera states for pretraining action-conditioned world models, solving the lack of control-signal supervision in real-world video.
- The pipeline runs in two stages: Stage I executes real-time physics in PIE, recording per-frame character states, control inputs, and camera states; Stage II replays trajectories in a fresh engine process for high-quality offline rendering via Movie Render Queue.
- A distributed system handles cache-aware task partitioning, node-local slot scheduling, automated scene screening, aesthetic/luminance filtering, partial-output recovery, async upload, and cluster health monitoring. The cluster spans 25 servers with 8 NVIDIA RTX 5090s each, processing 2,384 scene tasks.
More from Infra
- Intent-based multi-model routing: autonomous agents manage Sparks-hosted Qwen, GLM, DeepSeek — jasonkneen · 2026-09-06
- Data center spend jumps $25B in six months as construction hiring rebounds — a16z · 2026-09-06
- Musk: 3D printing enables integrated flow paths but is too slow and costly for volume production — i_bioloid · 2026-09-06
- The Economist: The moral panic over data centres is foolish — andsoitis · 2026-09-06
- Nvidia guides 70% revenue growth next year, supply sets the ceiling — BenBajarin · 2026-09-06
- PyPI's Recent Download Corrections Sharply Cut Some Packages' Stats — dbreunig · 2026-09-06