Stream3D: Generating Consistent 3D Objects from Streaming Video
机器之心 · wechat · 2026-08-02
A joint team from HKUST, MIT, and Harvard introduced Stream3D, combining 3D reconstruction's observational fidelity with 3D generation's prior completion. The framework adds a training-free streaming mechanism to frozen 3D generators (e.g., SAM3D, TRELLIS), enabling the creation of consistent 3D models from continuous monocular video streams.
The core is the Adaptive Evidential Memory:
- Attention Probes: Evaluates view reliability via cross-attention during a single denoising step.
- Fixed-Capacity Memory: Maintains a fixed-size list of reliable historical views per token, preventing memory bloat as video length increases.
- Multi-View Voting: Tokens collectively vote for the top-K most representative frames to condition the final generation.
Experiments show Stream3D outperforms frame-by-frame generation and pure streaming reconstruction on geometry and appearance metrics, effectively resolving temporal inconsistencies and missing surfaces in video-based 3D generation.
More from Multimodal
- OraRL: Efficient and Scalable RL for Video MLLMs — Yunheng Li · 2026-08-26
- Seeking Fast HD MiniMax Video Generation Without Quality Loss — OkMeat6773 · 2026-08-26
- Emotional animation of girl touching sky whale generated by Google Gemini — michaelrabone · 2026-08-26
- Wan 3.0 vs. MiniMax H3 Video Comparison: Smoother Audio and Transitions — SimplyAnnisa · 2026-08-26
- AI generated dance video shows cool moves — No-Bookkeeper-char · 2026-08-26
- Face Anything: 4D Face Reconstruction from Any Image Sequence (ECCV 2026) — rsasaki0109 · 2026-08-26