Stream3D: Training-Free Framework Enables Stable 3D Object Generation from Video Streams
青稞AI · wechat · 2026-07-30
A joint research team from HKUST, MIT, and Harvard introduced Stream3D, a novel framework that bridges 3D reconstruction and 3D generation to produce complete and consistent 3D objects from continuous monocular video streams.
Core Mechanism: Adaptive Evidential Memory
- Training-Free: The approach does not require modifying or retraining the underlying 3D generator (e.g., SAM3D, TRELLIS). Instead, it adds a training-free streaming mechanism.
- Constant Memory Size: Using attention probes, the model evaluates the reliability of the current perspective and retains only the historical frames with the highest "evidence scores" for each local region. The memory footprint remains constant and does not scale linearly with video length.
- Voting-based Generation: During the final generation phase, local regions vote for the historical frames that best explain them. The system selects the Top-K perspectives to jointly drive generation, effectively completing unobserved back structures.
Experimental Results
On the GSO and NAVI datasets, Stream3D outperforms baseline methods like frame-by-frame generation or latent state passing (e.g., KV-Cache) across multiple geometric and appearance metrics. It not only improves visual textures but also substantially enhances the recovery accuracy of 3D geometric structures, proving that generative models constrained by continuous evidence can achieve better completeness and consistency than pure streaming reconstruction.
Related event: Stream3D Enables Stable 3D Generation from Video Streams(2 posts)→
More from Multimodal
- OraRL: Efficient and Scalable RL for Video MLLMs — Yunheng Li · 2026-08-26
- Seeking Fast HD MiniMax Video Generation Without Quality Loss — OkMeat6773 · 2026-08-26
- Emotional animation of girl touching sky whale generated by Google Gemini — michaelrabone · 2026-08-26
- Wan 3.0 vs. MiniMax H3 Video Comparison: Smoother Audio and Transitions — SimplyAnnisa · 2026-08-26
- AI generated dance video shows cool moves — No-Bookkeeper-char · 2026-08-26
- Face Anything: 4D Face Reconstruction from Any Image Sequence (ECCV 2026) — rsasaki0109 · 2026-08-26