SemanTok: 49M autoregressive video model beats SOTA rivals 47x its size
CSProfKGD · x · 2026-10-06
SemanTok, a summer project at Stability AI, introduces predictable semantic tokens for autoregressive video generation. The 49M-parameter model beats SOTA VideoFlexTok — up to 47x larger — on semantics, and at 201M it wins in fidelity against a model 3.4x its size.
More from Multimodal
- AI-generated video is so funny netizens say the compute was worth it — hexiang · 2026-10-06
- Midjourney + Threejs + Seedance workflow for multi-angle AI cinematography — Ror_Fly · 2026-10-06
- Suno turns pictures into songs — and it even recognized the Gram-Schmidt meme — anderssandberg · 2026-10-06
- Apple researcher lands two NeurIPS 2026 papers on normalizing flows, including Normalizing Trajectory Models — thoma_gu · 2026-10-06
- Magnific teases October 8 launch, simple prompts already yield impressive results — aziz4ai · 2026-10-06
- Gaussian GRPO normalizes multimodal RL reward distributions, boosting OpenVLThinker v2 — kaiwei_chang · 2026-10-06