GigaChat Audio targets long-form audio grounding with timestamps across 120-minute inputs
ai-sage · hf · 2026-07-21
- GigaChat Audio is a time-aware large audio language model designed for temporal grounding in long recordings.
- It can answer questions with explicit timestamps over inputs as long as 120 minutes.
- The method interleaves periodic time markers with continuous audio tokens and uses large-scale synthetic supervision from a cascaded pipeline.
- Reported results show strong temporal-grounding accuracy on both short and long benchmarks, plus support for time-anchored fragment descriptions and summaries.
- The release also includes model weights and datasets for further research.
More from Research
- Microsoft open-sources Resource2Skill to turn videos and articles into executable agent skills — aigclink · 2026-07-21
- Microsoft open-sources Resource2Skill to turn tutorials into executable agent skills — aigclink · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21
- Shared agent workspaces fail in a fixed order, from stale reads to zombie writes — mrvladp · 2026-07-21
- A solo founder built an enterprise knowledge graph without a graph database — TheRedfather · 2026-07-21
- Practical rolling-shutter pose estimation uses affine correspondences — ducha_aiki · 2026-07-21