Alibaba open-sources TLive-Omni: omni-modal model for e-commerce live streaming with 256K context

aigclink · x · 2026-08-26

Alibaba Taobao has open-sourced TLive-Omni, an omni-modal understanding model for e-commerce live streaming. Built on a Qwen3.5 backbone with an AuT audio encoder, it supports 256K token context. The model uses a Timestamped Per-vGrid layout to align audio and video tokens by time grid, solving second-level audio-visual alignment. It supports speech recognition, speaker identification, product visual grounding, OCR, temporal localization, dense video captioning, omni-modal QA, and multi-dimensional shot annotation. Versions 4B and 9B are available, and the technical report is on arXiv. Applications include building live replay systems, review tools, and content clipping.

Related event: TLive-Omni: Open-Source Omni-Modal Model for Livestream E-commerce(4 posts)→

Original post →

More from Models

Models channel →