Fourier deformation forces a 10s loop on static 3DGS, preferred in up to 99% of votes

OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs

You-Zhe Xie, Ting-Wei Chou, Yu-Hsuan Li, Kaipeng Zhang, Zhixiang Wang, Yu-Lun Liu

cs.CV, cs.GR

2026-10-09

OuroWorld turns any static 3DGS into an endless free-viewpoint loop by a Fourier deformation. On 39 scenes, users prefer it in 70.8% to 99.0% of votes.

What problem this solves

World models such as HY-World 2.0, Marble, and Lyra 2.0 already export walkable, photoreal scenes. Leaves, mist, and lights stay still. A 2D cinemagraph adds a seamless local loop to a still. In 3D, LoopGaussian and 3D-MOM drive a 3D Gaussian Splatting scene, a cloud of Gaussian primitives, with an Eulerian flow: a velocity field on a fixed grid. That prior fits smoke, water, and fog. It does not fit general deformation, object displacement, or lighting change. Layered depth images break once the camera moves, and every prior method also needs a motion mask.

Video generators can supply diverse, roughly periodic motion. Two gaps remain. The clip is only approximately periodic, so a fitted loop jumps at the seam. One view cannot determine a 4D scene, and generated side views disagree, so a direct fit blurs the geometry and makes the motion jitter.

OuroWorld takes any static 3DGS and decides what should move. Inside the tested horizontal orbit of ±20°, motion stays consistent across views, and the representation itself returns to the start of the cycle. On the paper's capability checklist it is the only entry that is mask-free, truly 3D, periodic by construction, not limited to fluid motion, and able to change illumination.

Method

Three stages. The input is a static 3DGS.

GPT-5.5 reads a rendered reference view and writes a motion prompt: visible motion, several elements when the scene calls for it, no new objects, a fixed camera, and a natural loop. Seedance 2.0 uses that image as both the first and last frame and generates a 10-second, 720p, 24 FPS clip. Unnatural clips are filtered out. This replaces the mask.

VGGT-Ω reads 24 frames of that clip plus 21 static orbit renders at t=0. The stills only steady the depth estimate. The model lifts a per-frame point cloud. Twenty cameras sit within ±20° of a hand-picked pivot. The cloud is rendered with holes, and TrajectoryCrafter fills them conditioned on the reference clip. Those views share one initial noise.

The scene model is Inconsistency-Robust Periodic 4DGS. A deformation field queries a triplane, three axis-aligned feature planes, at each Gaussian center and splits the feature into Fourier coefficients: a constant term plus K=4 sine and cosine pairs, with period T=10 seconds. An MLP of width 128 maps that feature to offsets on position, rotation, scale, and color. After one period the sines and cosines repeat, so time t and time t+T are the same scene however optimization finishes. Earlier Fourier 4D models used the basis for capacity. Here it is what forces the loop.

Fitting the side views directly makes the shared field average contradictory observations, and the scene goes soft. A drift field on every view, including the reference, can explain the motion by itself. At inference that drift is discarded and the scene goes quiet. The grounded variant sets drift to zero on the reference view, so the deformation field has to reproduce that clip alone. Drift on the other views only absorbs their residual. Inference renders Gaussians plus deformation.

The reference video and the input scene's t=0 views get an L1 loss. Other generated frames are not pixel targets. The current render is noised and denoised with Stable Diffusion v1.5 while the latent is pulled toward the generated frame, the SDEdit pattern, and LPIPS then matches the render to that refined image. The loss weights are 1 and 0.2. Every run uses one RTX 4090.

Results

The benchmark has 39 scenes: 9 reconstructed from Mip-NeRF 360, and 10 each from the three world models. Renders last 20 seconds at 30 FPS and 672×384. One camera stays fixed. The other yaws through 0°, 20°, -20°, and back to 0°, at 0.95 times the training radius. Native periods are repeated rather than retimed: 0.5 s, 2.4 s, 2.0 s, 2.0 s, and 10 s for Gaussians-to-Life, 3D Cinemagraphy, 3D-MOM, LoopGaussian, and OuroWorld. Baseline masks come from the same dynamics prompt, SAM 3, and a human check.

Vividness averages motion, illumination, and appearance change at 3 FPS. A pixel counts as moving when optical flow exceeds 0.3% of the short side, and as relit when luminance shifts by more than 20%. KVD is measured against 780 clips from MiniMax H3, not Seedance. MALF and Seam SSIM are static-camera scores. MALF credits a return to the same phase one period later and is 0 for a still. Seam SSIM compares only the two frames across the cut.

MethodVividness ↑KVD ↓MALF ↑Seam SSIM ↑
Gaussians-to-Life0.0374127.470.02650.9388
3D Cinemagraphy0.0066142.420.01340.9995
LoopGaussian0.0024154.650.00250.9998
3D-MOM0.0247155.230.03170.9158
OuroWorld0.0423108.230.08090.9965

On the fixed camera, 0.0423 versus 0.0374 is a small vividness lead over Gaussians-to-Life. Raw motion variation goes the other way: MV 0.0588 versus 0.0413. The 0.5-second period inflates frame-to-frame change, while MALF stays at 0.0265. Illumination is where OuroWorld pulls away: IV 0.0518 against 0.0248 for 3D-MOM and 0.0170 for Gaussians-to-Life. Seam SSIM is 0.9965, below LoopGaussian at 0.9998 and 3D Cinemagraphy at 0.9995. KVD is 108.23 against a next-best 127.47.

Under the orbit, camera motion dominates the frame. Vividness is 0.5117 versus 0.5097, and KVD is 50.11 versus 51.86. Aesthetic Quality is best at both cameras.

Thirty people ran a forced choice: 780 pairs, five judges each. Preference for OuroWorld spans 70.8% to 99.0%. Against 3D-MOM, vividness and the seam both get 99.0%. Against LoopGaussian, naturalness and visual quality get 70.8%, so 29.2% of judges picked the near-still side. People separate the methods more than the orbit metrics do.

Removing the grounding drops static vividness from 0.0423 to 0.0392 and MV from 0.0413 to 0.0324. Setting the Fourier period to 2T drops Seam SSIM from 0.9965 to 0.9783, while MALF only moves from 0.0809 to 0.0788. Dropping the diffusion refinement raises static vividness to 0.0450. The appendix reads that rise as flicker from inpainting error baked into the scene. Orbit sharpness, VoL, falls from 709.75 to 551.33. Removing the drift field alone leaves vividness at 0.0424.

Why it matters

Given a static 3DGS or a world-model export, this is a mask-free post process. Deformation, displacement, and lighting changes share one representation, and the orbit still closes. Inference drops the drift field, so runtime is Gaussian rasterization plus a small MLP.

Against the Eulerian baselines, a wider motion range and a hard loop arrive together. Against Gaussians-to-Life, which also uses a video prior, static vividness only moves from 0.0374 to 0.0423, while MALF moves from 0.0265 to 0.0809 and illumination from 0.0170 to 0.0518. The gain is a viewable, closed 3D loop, not a sudden jump in how hard things move.

Limitations

The video prior caps the dynamics. A bad generation survives into the scene, and VLM filtering only makes that less common. Every side view is completed from one reference clip, so a wider orbit means more disocclusion and more inconsistency. The tested path is ±20°, narrower than the abstract's claim of arbitrary viewpoints.

There is no ground truth. Vividness thresholds were tuned by eye. KVD treats generated video as the natural distribution, and OuroWorld is distilled from a video model, so Eulerian baselines are handicapped on that score. MiniMax H3 avoids Seedance. It does not remove the bias. Scenes were chosen for water, plants, cloth, fire, or lights. Pivots were hand-picked, and masks were checked by hand. Thirty-nine scenes do not support a claim about every export. Periods from 0.5 to 10 seconds share one 20-second timeline, so the methods are not compared at the same speed. Replacing GPT-5.5 and Seedance 2.0 is untested. The drift field's sharpness gain is smaller than the diagram suggests: without it, static vividness is 0.0424 against 0.0423, and the loss of detail over time shows up mainly in the qualitative examples.

Terms

Source

Related papers

All paper explainers