DynaTokens: EMNLP paper generates adaptation tokens on demand for video-language learning

flosalim · x · 2026-09-22

flosalim announces DynaTokens (Controlling Token Dynamics for Continual Video-Language Understanding), accepted to EMNLP 2026 Main. It tackles continual adaptation in video-language models: instead of storing a separate set of adaptation tokens per task, the model learns to generate the right tokens on demand. Built on a frozen LLaMA-2-7B backbone with CLIP ViT-L/14 and LLaMA-Adapter for efficient adaptation, no backbone fine-tuning needed.

Related event: DynaTokens Accepted to EMNLP 2026 for Continual Video-Language Adaptation(2 posts)→

Original post →

More from Multimodal

Multimodal channel →