TwelveLabs shows a video memory layer that stores moments, entities and themes

AI Engineer · youtube · 2026-07-23

In a talk titled Video Has No Memory. Here's How We Built One., TwelveLabs’ James Le argues that most video AI systems answer each query from scratch because video lacks a durable memory representation.

He says bigger context windows do not solve the core problem. Instead, the right approach is to treat video as a spatial-temporal volume and build a memory layer on top of it:

TwelveLabs’ stack includes an embedding encoder, a context store, and a video language model exposed as an API. The demo examples span World Cup highlights, tracking Messi across a corpus, traffic security, and ad placement.

Original post →

More from Multimodal

Multimodal channel →