Building Training Supervision Offline Using 40k Videos
mangahomanga · x · 2026-07-18
The author details the sources of their training data and supervision signals:
- Offline processing of 40,000 Something-Something-V2 interaction videos to extract object masks, metric geometry, camera motion, and dense 3D tracks.
- Future frames are used strictly as supervision targets, not as model inputs.
- This means the model learns to predict future motion based solely on past observations, without relying on future information leakage.
Related event: MotionForesight predicts future 3D motion from existing video models(6 posts)→
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Turning Noise into Signal: Predicting TCR Binding Using AlphaFold3 Hallucinations — quaidmorris · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22