Surflo (NeurIPS 2026 Oral): 3–60 unposed photos into one coherent 3D surface via flow matching
RexDouglass · x · 2026-09-25
Researchers from École Polytechnique, Kyoto University, Kyutai, and UC Berkeley present Surflo, accepted as a NeurIPS 2026 Oral.
- Encodes a variable number of unposed views (3–60) into a single fixed-size global latent (K=128 tokens, VGGT backbone + Perceiver compressor), unlike per-view pointmap models (VGGT, DUSt3R, DepthAnything-3) whose representations grow linearly with views, adding noise and redundancy.
- Each surface point is decoded independently with a flow-matching ODE conditioned on the latent—any number of points (up to 10^6) and resolution from one encoder pass. A guidance mechanism with shared rendering loss (points rendered as 3D Gaussians) couples nearby points for coherent surfaces.
- State of the art on 8 benchmarks (2–32 views); releases a new real-world surface dataset of 10.5K DL3DV scenes with full meshes. Code, data, and arXiv are public.
More from Research
- MUX latent reasoning method wins NeurIPS Spotlight, cuts CoT length 3-6x — mmbronstein · 2026-09-25
- Stanford's 37K AI agents analyzed 57K clinical trials as a 'virtual biotech' — james_y_zou · 2026-09-25
- CVP from UC San Diego and Lambda beats Video-3D-LLM on all five 3D spatial reasoning benchmarks — TheZachMueller · 2026-09-25
- NeurIPS 2026 decisions for main, Position and E&D tracks out today on OpenReview — NeurIPSConf · 2026-09-25
- Auto-Research Arena: 6,300 runs, agents rediscover MQA/MLA, memory layers and more — qixing_huang · 2026-09-25
- Economists debate AI research: clean identification takes decades, but we need signals now — daveholtz · 2026-09-25