Stability AI's SemanTok: 201M AR video model matches a 3.4x larger rival with semantic tokens
stabilityai · hf · 2026-10-02
Stability AI published SemanTok, a flexible-length semantic video tokenizer for autoregressive video world models.
Core idea: feed frozen DINO features into the tokenizer encoder and add lightweight heads that reconstruct them from each retained token prefix, achieving strong semantic alignment at every noise level (unlike REPA-only approaches).
Results:
- A 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4x its size
- Coarse-to-fine prefixes carry global semantics first, pixel detail deferred to later tokens — making short prefixes cheaper to predict with better fidelity
- Holds semantic alignment on out-of-distribution classes; works well for both reconstruction and generation
More from Multimodal
- Full music video generated locally with ComfyUI and LTX 2.5 on 16GB VRAM — sokmech · 2026-10-02
- Eleven v4 character animation demo shown off in new video — buraktuyan · 2026-10-02
- Claude directs its own EDM music video 'The Good Ending' in Fable 5.1 demo — cheetoskull · 2026-10-02
- All-in-one Krea 2 Turbo ComfyUI workflow uses 2/4-step distilled LoRAs for speed — TimeTruth2490 · 2026-10-02
- No one on the team knew Blender — Claude ran the whole 3D film workflow — gen_ericai · 2026-10-02
- Meshy teases Meshy Edit: tweak specific parts of 3D models with a text prompt — rms80 · 2026-10-02