Alibaba open-sources TLive-Omni, an omni-modal model for e-commerce live streaming with 256K context

aigclink · x · 2026-08-26

Alibaba Taobao has open-sourced TLive-Omni, an omni-modal understanding model for e-commerce live streaming, capable of processing video, audio, and text simultaneously. Built on a Qwen3.5 backbone with an AuT audio encoder, it supports 256K token context, enabling hours of live content memory. The core innovation is a Timestamped Per-vGrid token layout that aligns audio and video tokens by time grid, solving the problem of second-level audio-visual alignment. The model supports speech recognition, speaker identification, product visual grounding, OCR, temporal localization, dense video captioning, omni-modal QA, and multi-dimensional shot annotation. Applications include building live replay systems, review tools for streamers, and content clipping.

Related event: TLive-Omni: Open-Source Omni-Modal Model for Livestream E-commerce(4 posts)→

Original post →

More from Models

Models channel →