DynaTokens: EMNLP paper generates adaptation tokens on demand for video-language learning
flosalim · x · 2026-09-22
flosalim announces DynaTokens (Controlling Token Dynamics for Continual Video-Language Understanding), accepted to EMNLP 2026 Main. It tackles continual adaptation in video-language models: instead of storing a separate set of adaptation tokens per task, the model learns to generate the right tokens on demand. Built on a frozen LLaMA-2-7B backbone with CLIP ViT-L/14 and LLaMA-Adapter for efficient adaptation, no backbone fine-tuning needed.
Related event: DynaTokens Accepted to EMNLP 2026 for Continual Video-Language Adaptation(2 posts)→
More from Multimodal
- MiniMax H3 video gen runs locally on M5 Ultra: 768p in ~2m22s with optimizations — bakawolf123 · 2026-09-22
- Creator Gives Away His Entire Prompt Collection, Organized by Scene in a Milanote Canvas — gen_ericai · 2026-09-22
- ip-as-logo-skill: open Agent Skill that generates cute rounded IP mascot logos, 5.4k stars — tom_doerr · 2026-09-22
- Seeking a local/free workflow for 5-shot character-consistent video (Krea + H3 + ComfyUI) — NewPhoneWhotiz · 2026-09-22
- Qwen 2.1 CFG test: CFG 1.0 + 12 steps halves image generation time — FreeTheClanks · 2026-09-22
- Qwen-Image-2.1 Tops Open Text-to-Image Arena at 1228 pts, Just 11 Behind GPT-Image — arena · 2026-09-22