Alibaba open-sources TLive-Omni, an omni-modal model for e-commerce live streaming with 256K context
aigclink · x · 2026-08-26
Alibaba Taobao has open-sourced TLive-Omni, an omni-modal understanding model for e-commerce live streaming, capable of processing video, audio, and text simultaneously. Built on a Qwen3.5 backbone with an AuT audio encoder, it supports 256K token context, enabling hours of live content memory. The core innovation is a Timestamped Per-vGrid token layout that aligns audio and video tokens by time grid, solving the problem of second-level audio-visual alignment. The model supports speech recognition, speaker identification, product visual grounding, OCR, temporal localization, dense video captioning, omni-modal QA, and multi-dimensional shot annotation. Applications include building live replay systems, review tools for streamers, and content clipping.
Related event: TLive-Omni: Open-Source Omni-Modal Model for Livestream E-commerce(4 posts)→
More from Models
- Zhipu GLM-5.3-Flash: Matches Opus 4.8 at 1/40 the Cost, Powered by Domestic Chips — vista8 · 2026-08-27
- TokenSpeed adds Day-0 support for Qwen 3.8 Flash Next architecture — Alibaba_Qwen · 2026-08-27
- Zhipu GLM-5.3 open weights releasing in 22 hours — Yuchenj_UW · 2026-08-27
- AI models show more creativity when talking to each other than in assistant persona — nabeelqu · 2026-08-27
- OpenRouter leaderboard: Real token consumption data outweighs media hype — sujingshen · 2026-08-27
- Qwen 3.8-Next Released with Detailed Technical Report on Architecture — nrehiew_ · 2026-08-27