Stability AI's SemanTok: 201M AR video model matches a 3.4x larger rival with semantic tokens

stabilityai · hf · 2026-10-02

Stability AI published SemanTok, a flexible-length semantic video tokenizer for autoregressive video world models.

Core idea: feed frozen DINO features into the tokenizer encoder and add lightweight heads that reconstruct them from each retained token prefix, achieving strong semantic alignment at every noise level (unlike REPA-only approaches).

Results:

Original post →

More from Multimodal

Multimodal channel →