Direct 4D Training Hits 1T Context: Pixel/Point Cost Comparison in Video, 3D, and Interactive Gen
srchvrs · x · 2026-08-26
A discussion on context window scale references data on "direct 4D training" at 1T context with inference past 5T. It compares raw pixel/grid point counts across modalities:
- Video Gen (Veo 3.1, Kling 3.0): 1B per clip (8–15s, 1080p–4K), generated in one shot without explicit 3D.
- Marble: Static 3D built from 10⁷ px panorama, lifted into a persistent scene of roughly as many points.
- Genie 3: Interactive 720p @ 24fps, generated frame by frame, with 1.3B pixels per minute of memory.
More from Research
- CAFE Improves Search Agents via Co-Evolving Agent and Critic — _reachsumit · 2026-08-26
- OPDSearch+: distill a frozen teacher then refine with RL to beat RL-only training — _reachsumit · 2026-08-26
- Retrieval Grounding Outperforms Confidence Voting for Search Agents — _reachsumit · 2026-08-26
- AdaWidth Adapts Embedding Dimensions for Efficient Dense Retrieval — _reachsumit · 2026-08-26
- Speculative Programmatic Tool Calling Speeds Up Code Agents by 1.2x — a1zhang · 2026-08-26
- CodeHID: Generative Code Retrieval via Hierarchical Indexing — _reachsumit · 2026-08-26