Sori-1B: Audio-Grounded LM Trained From Scratch With No Text-Only Pretraining
Balance- · reddit · 2026-08-31
An SNU researcher released Sori-1B, a 1B-parameter audio-language model whose decoder is trained entirely from scratch on audio-paired text, with no text-only pretraining or pretrained-LM initialization.
Technical Highlights:
- Audio-Grounded: Designed to ground answers in audio rather than relying on text-only priors (experiments show AF3 retains 74% performance with silence, indicating text reliance).
- Architecture: Reuses NVIDIA's frozen Audio Flamingo Next encoder (61.5% of params); everything else (decoder, embeddings, custom audio-concept tokenizer) is trained from scratch.
- Training Scale: Trained on 7.4k hours / 4.75M examples using just 3x RTX 4090s.
- Capabilities: Supports MCQ, open QA, captioning, and ASR.
- License: Weights are gated for non-commercial/academic use only due to NVIDIA's encoder terms.
More from Multimodal
- French Director Releases AI-Generated Short Film 'La Nona Gigante' — venturetwins · 2026-08-31
- MiniMax H3 Max Generates Video Faster Than Playback Speed — isidentical · 2026-08-31
- Turning mental rabbit holes into moving images: A generative video experiment — Kyrannio · 2026-08-31
- Grok vs GPT-4o Image: Striking alignment revealed by same prompt — teortaxesTex · 2026-08-31
- One Prompt, One Documentary: H3 Generates Hand-Drawn Video Explaining Espresso vs Americano vs Cappuccino — Best_Candidate_9060 · 2026-08-31
- Running MiniMax H3 on an 8GB laptop: full I2V pipeline with configs and failures — DG86 · 2026-08-31