Meta AI: Mid-training LLM on raw video lifts multimodal scores without text loss
kastnerkyle · x · 2026-10-10
A Meta AI paper (arXiv:2610.11019) tests whether raw, caption-free web video can serve as mid-training data for a pretrained LLM.
- Method: video frames are encoded into continuous visual tokens; Qwen3-1.7B is mid-trained via next-visual-token prediction on YT-Temporal-1B, fully self-supervised with no text loss
- Results: +2.9 points average across 4 video benchmarks and +5.1 across 10 image benchmarks vs. no mid-training; text performance preserved (48.9 vs 48.0 over 14 text benchmarks)
- Training dynamics: image/video gains emerge within the first 30% of training then plateau (within 0.5 points)
- Predicting captions underperforms next-visual-token prediction, avoiding captioning compute and labeling noise
More from Multimodal
- AI fake videos have hit another level of realism, researcher warns — rohanpaul_ai · 2026-10-10
- Designer ships polished launch video entirely in code: Remotion, Three.js and Suno, no editor — lmoroney · 2026-10-10
- AI-generated trailer wins inaugural $2.5 million Future Vision XPRIZE — DavidmComfort · 2026-10-10
- A beginner-friendly ComfyUI/MiniMax terminology cheat sheet from Reddit — wildmonkeywrangler · 2026-10-10
- HiDream-O1-Video-1.0 debuts #6 on image-to-video leaderboard at $5.80/min — ArtificialAnlys · 2026-10-10
- AI-cast drama series 'The Engagement' drops episode 4 on Rad TV — AIandDesign · 2026-10-10