GigaChat Audio 10B adds speech understanding with a Conformer and MoE decoder

pmttyji · reddit · 2026-07-26

GigaChat Audio 10B is an audio-native LLM built on top of GigaChat 3.1 Lightning. It uses a Conformer speech encoder plus a modality adapter to feed audio embeddings into a Mixture-of-Experts decoder, preserving the base text model’s quality while adding speech understanding.

The model supports audio QA, classification, temporal grounding with timestamps, audio summarization, tool use, and text-only tasks. Its grounding ability was trained on TimeGround-1M, a long-form audio dataset with time-aligned annotations.

Original post →

More from Multimodal

Multimodal channel →