SemanTok: 49M autoregressive video model beats SOTA rivals 47x its size

CSProfKGD · x · 2026-10-06

SemanTok, a summer project at Stability AI, introduces predictable semantic tokens for autoregressive video generation. The 49M-parameter model beats SOTA VideoFlexTok — up to 47x larger — on semantics, and at 201M it wins in fidelity against a model 3.4x its size.

Original post →

More from Multimodal

Multimodal channel →