RegVGGT: training-free token regulation keeps only 1% of tokens per frame for streaming 3D reconstruction
zhenjun_zhao · x · 2026-09-22
An ECCV 2026 paper from Harbin Institute of Technology tackles the memory dilemma of feed-forward reconstruction models on long video streams.
- Key observation: a token's initial saliency reliably dictates its long-term importance across the stream
- Method: a training-free token regulation scheme admits at most 1% of tokens per frame to update context memory, with a FlashAttention-compatible saliency estimator
- Results: thousands of frames processed on a consumer GPU with negligible reconstruction loss, setting SOTA on long-horizon benchmarks
More from Research
- Rich RL Report Ships With 9B Distilled Model and 7,000 Open RL Environments — tokenbender · 2026-09-22
- Reading Xiaomi's MoE RL Release: Specialist Teachers Beat Big Mixed RL on Hard Tasks — tokenbender · 2026-09-22
- New arXiv Paper: Router-Aware Importance Sampling Stabilizes MoE RL Training — tokenbender · 2026-09-22
- Sticking with GRPO: The Four Targeted Mods Behind Xiaomi's Stable MoE RL — tokenbender · 2026-09-22
- RL training stability tricks: prompt-mean loss, asymmetric clipping and entropy-adaptive GRPO — tokenbender · 2026-09-22
- Thermofluids professor: OpenAI's Clay problem solution is not physically reproducible — GaryMarcus · 2026-09-22