Making 370 Hours of Archival Video Searchable with VLMs for $10
vanstriendaniel · x · 2026-07-31
The author shared a practical project using Vision-Language Models (VLMs) to make video archives searchable by on-screen actions. He processed 1,864 public-domain films from the Prelinger Archives, cutting them into 60-second chunks and indexing them using an open-source 2B parameter video model.
Building the entire index for 370 hours of video cost only about $10 in GPU time. The pipeline allows for highly specific semantic searches, such as querying "typing on a computer keyboard" to jump directly to the exact moment in the footage.
Related event: AI Makes 370 Hours of Public Domain Films Searchable for $10(3 posts)→
More from coding & agent
- Free Lesson: Turn Eval Results Into a Better AI Model — xeophon · 2026-07-31
- NVIDIA Tutorial: Run Your Polars Data Processing Code on the GPU — NVIDIAAI · 2026-07-31
- Open-Source AI Agent Validates Startup Ideas in 10 Minutes — tom_doerr · 2026-07-31
- Open-Sourcing Agents: Code is 10%, Surviving Unknown Infra is 90% — Future_AGI · 2026-07-31
- Positron's Jupyter Notebook Editor Goes GA: AI-Powered IDE for Data Science — HamelHusain · 2026-07-31
- System prompts are constraint specs, not descriptions: a 4-step framework to stop LLM degradation in production — ClickOk5811 · 2026-07-31