Qwen2-VL handles 11-min video with 96 timestamped events accurately

vanstriendaniel · x · 2026-08-15

A developer tested Qwen2.5-VL on an 11-minute film from 1935 in a single request. The model generated 96 timestamped events within 157 seconds, quoting on-screen text verbatim.

Manual verification showed timestamps were accurate to within 2 seconds of the extracted frames. The test ran on a single GPU via Hugging Face Jobs. The author has open-sourced a PoC script for whole-video captioning and timeline extraction without chunking.

Related event: Qwen Model Accurately Analyzes 11-Minute Long Video(2 posts)→

Original post →

More from Multimodal

Multimodal channel →