GigaChat Audio 10B adds speech understanding with a Conformer and MoE decoder
pmttyji · reddit · 2026-07-26
GigaChat Audio 10B is an audio-native LLM built on top of GigaChat 3.1 Lightning. It uses a Conformer speech encoder plus a modality adapter to feed audio embeddings into a Mixture-of-Experts decoder, preserving the base text model’s quality while adding speech understanding.
The model supports audio QA, classification, temporal grounding with timestamps, audio summarization, tool use, and text-only tasks. Its grounding ability was trained on TimeGround-1M, a long-form audio dataset with time-aligned annotations.
More from Multimodal
- Krea AI creates, edits, and upscales images and videos in real time — Med1_Ai · 2026-07-26
- Ideogram generates high-quality AI images with accurate text rendering — Med1_Ai · 2026-07-26
- A reusable blurred-silhouette prompt for motion-heavy sports images — azed_ai · 2026-07-26
- Seedance 2.0 is shown with a cinéma vérité video prompt — techhalla · 2026-07-26
- A smartphone clip becomes an audiovisual piece for under 50 cents in a mocap alternative — uisato · 2026-07-26
- ComfyUI tutorial shows how to remove image backgrounds with SAM3 segmentation — Main-Strawberry9241 · 2026-07-26