ATGS: Anchored Temporal Gaussian Splatting for Long Volumetric Video Representation
Jiahao Wu, Jie Liang, Die Hu, Jiayu Yang, Kaiqiang Xiong, Xiang Li, Xiaoyun Zheng, Chao Wang, Ronggang Wang
SIGGRAPH'2026)
cs.CV
2026-08-31
ATGS localizes Gaussians around time-conditioned anchors and a width-7 window, reconstructing 1,400-frame fast-motion video in one run; VRU Long PSNR 24.78 vs LocalDyGS 23.21.
Volumetric video turns multi-view 2D footage into a 4D scene you can inspect from any camera and any timestamp. Current 3D Gaussian Splatting and 4D Gaussian pipelines keep failing at two demands that pull opposite ways: minute-scale sequences with thousands of frames, and fast motion with large displacements, several people, and lighting change.
LongVolCap can run at minute length, then smears when people move quickly. FreeTimeGS and LocalDyGS track complex motion, but only for one or two seconds. Chopping a long take into clips and stitching them produces flicker at the cuts. Asking one Gaussian primitive to follow a long, messy trajectory is an unstable optimization.
ATGS's bet is concrete: stop asking points to travel far. A sparse set of time-stamped anchors localizes where Gaussians may appear.
Time-conditioned anchors are initialized from periodically sampled keyframes. Each stores a 3D position, a 64-d feature fa, a spatial scale v, and a keyframe index k. Anchors control when Gaussians appear and vanish. They do not fit explicit point trajectories.
At query time t, only anchors around the nearest keyframe stay active. The default window width is W=7. Distant timestamps stay out of the decode, so a long sequence becomes a stack of local jobs.
Three features then constrain the decoder. fa holds local detail and is optimized per anchor. fs is trilinearly sampled from a shared 128³×64 spatial grid and supplies time-invariant local consistency, so neighboring anchors do not drift apart. ft comes from 4D hash grids sliced into M temporal segments: M=2 for indoor small motion, M=12 for VRU basketball. Each query uses only the local time inside its segment.
The decoder concatenates [fa+fs; ft] and runs a two-layer MLP to emit mean, orientation, scale, opacity, and color. Concatenation, rather than a sum, keeps static and dynamic parts from having to disentangle themselves; summing them slows training by about 10%. The loss is L1 plus SSIM at weight 0.2, plus a volume regularizer at 0.001 that keeps Gaussians compact around their anchor.
Small-motion scenes (N3DV, MeetRoom) use K=15 keyframes. VRU basketball uses K=250. Training runs on an A100 80GB GPU.
The stress test is VRU basketball: 36 cameras, GZ and DG at 250 frames each, Long at 1,400 frames of 4K fast play, about one minute. LocalDyGS-style methods usually train on about 20 frames. ATGS trains the full sequence in one run.
| Method | VRU GZ PSNR | VRU Long PSNR |
| 4DGaussian (short clips) | 28.32 | 23.01 |
| LocalDyGS (long) | 29.23 | 23.21 |
| ATGS (long) | 30.61 | 24.78 |
On VRU Long, ATGS sits 1.57 dB above LocalDyGS and 1.77 dB above 4DGaussian. On GZ it reaches 30.61, next to per-frame 3DGS at 30.50, with SSIM 0.948 versus 0.949.
On the 300-frame N3DV split, ATGS scores PSNR 32.56 against LocalDyGS 32.28 and SpaceTimeGS 32.05. Render speed is 70 FPS and training takes 0.9 h, slower than LocalDyGS at 105 FPS / 0.58 h, but one run covers 1,200 frames (about 40 s) where most baselines stop at 300. On sparse-view MeetRoom, PSNR is 32.79 against 3DGStream at 30.79, with a 110 MB model. On SelfCap Bar (2,000 frames) the score is 29.13 versus 4DGaussian 27.86.
Ablations pin the parts down. Dropping the temporal window on VRU Long moves PSNR from 24.42 to 24.02. Dropping fa and keeping only fs+ft collapses to 22.58. Raising temporal grids from M=1 to 12 lifts GZ PSNR from 29.10 to 30.61. Feature visualizations show fa plus fs already encode almost all scene content; ft mainly gates Gaussian appearance and disappearance, and does not learn a full dynamic geometry.
Volumetric-video engineering has been choosing between short-and-hard and long-and-easy. ATGS is a recipe you can retell: keyframe anchors, a fixed window, and a split between a static spatial grid and a dynamic hash. Anyone already on Scaffold-GS or LocalDyGS is looking at a temporalization of the same anchor-decode idea, not a new stack.
The usable setting is concrete: indoor conversation, a basketball court, a few minutes of a phone-array capture. One training pass over a thousand frames removes the clip-and-align post-process. On VRU-GZ, inference is 15.6 ms end to end (0.7 feature, 5.9 MLP, 4.4 rasterize), so playback is real-time. Capture and training stay offline.
On N3DV-style small motion the gain over LocalDyGS is 0.28 dB. The gap shows up on minute-scale fast action. If the product is a 300-frame indoor scene, this is incremental. If it is a one-minute court, that gap is the point.
The authors list three. The pipeline is offline reconstruction, not live capture. Poses and sparse points still come from COLMAP, the usual 3DGS tax. Distant, poorly observed regions such as the VRU stands show mild temporal jitter, because multi-view constraints are thin and moving subjects occupy few pixels.
The comparison table is also uneven. In Table 3, 3DGS, 4DGaussian, and SpacetimeGS are marked as short-clip training (20 frames or a single frame); only ATGS and LocalDyGS train long. Short-clip 3DGS nearly ties ATGS on GZ PSNR, so stretching the sequence has a cost. ATGS's claim is that it drops less on long takes, not that it beats short-clip overfitting. SelfVolCap and FreeTimeGS had no public code at the time, so they are absent from the tables. PKU-DyMVHumans is cited as a generalization test, with numbers left in the supplementary video.
Hyperparameters follow the scene: K=15 and M=2 for small motion, K=250 and M=12 for large. A new capture still needs a hand-chosen keyframe density and temporal-grid count. No selection rule is given.