TwelveLabs shows a video memory layer that stores moments, entities and themes
AI Engineer · youtube · 2026-07-23
In a talk titled Video Has No Memory. Here's How We Built One., TwelveLabs’ James Le argues that most video AI systems answer each query from scratch because video lacks a durable memory representation.
He says bigger context windows do not solve the core problem. Instead, the right approach is to treat video as a spatial-temporal volume and build a memory layer on top of it:
- ingest once, reason many times
- store primitives, not final answers
- ground every claim to timestamps
- let intent decide what gets remembered
TwelveLabs’ stack includes an embedding encoder, a context store, and a video language model exposed as an API. The demo examples span World Cup highlights, tracking Messi across a corpus, traffic security, and ad placement.
More from Multimodal
- AI lip-sync still breaks on pre-existing clips, even with matching audio — Kyrannio · 2026-07-23
- Reddit user shares an img2vid workflow built on an RTX 4070 Super — Professional_Wash169 · 2026-07-23
- NoSpoon music-video automation still breaks because models cannot really hear music — Kyrannio · 2026-07-23
- Early Krea 2 training results show the model already producing visuals — darlens13 · 2026-07-23
- ChatGPT Image 2.0 turns out a moody, classical-style painting — DeryaTR_ · 2026-07-23
- ChatGPT Work turns hundreds of Slack photos into a montage during a commute — gabrielchua · 2026-07-23