Voxtral: Mistral's 24B open audio chat model beats GPT-4o mini on transcription

Voxtral

Alexander H. Liu, Andy Ehrenberg, Andy Lo, Clément Denoix, Corentin Barreau, Guillaume Lample, Jean-Malo Delignon, Khyathi Raghavi Chandu, Patrick von Platen, Pavankumar Reddy Muddireddy, Sanchit Gandhi, Soham Ghosh, Srijan Mishra, Thomas Foubert, Abhinav Rastogi, Adam Yang, Albert Q. Jiang, Alexandre Sablayrolles, Amélie Héliou, Amélie Martin, Anmol Agarwal, Antoine Roux, Arthur Darcet, Arthur Mensch, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clémence Lanfranchi, Darius Dabert, Devendra Singh Chaplot, Devon Mizelle, Diego de las Casas, Elliot Chane-Sane, Emilien Fugier, Emma Bou Hanna, Gabrielle Berrada, Gauthier Delerce, Gauthier Guinet, Georgii Novikov, Guillaume Martin, Himanshu Jaju, Jan Ludziejewski, Jason Rute, Jean-Hadrien Chabran, Jessica Chudnovsky, Joachim Studnia, Joep Barmentlo, Jonas Amar, Josselin Somerville Roberts, Julien Denize, Karan Saxena, Karmesh Yadav, Kartik Khandelwal, Kush Jain, Lélio Renard Lavaud, Léonard Blier, Lingxiao Zhao, Louis Martin, Lucile Saulnier, Luyu Gao, Marie Pellat, Mathilde Guillaumin, Mathis Felardos, Matthieu Dinot, Maxime Darrin, Maximilian Augustin, Mickaël Seznec, Neha Gupta, Nikhil Raghuraman, Olivier Duchenne, Patricia Wang, Patryk Saffer, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Philomène Chagniot, Pierre Stock, Pravesh Agrawal, Rémi Delacourt, Romain Sauvestre, Roman Soletskyi, Sagar Vaze, Sandeep Subramanian, Saurabh Garg, Shashwat Dalal, Siddharth Gandhi, Sumukh Aithal, Szymon Antoniak, Teven Le Scao, Thibault Schueller, Thibaut Lavril, Thomas Robert, Thomas Wang, Timothée Lacroix, Tom Bewley, Valeriia Nemychnikova, Victor Paltz, Virgile Richard, Wen-Ding Li, William Marshall, Xuanyu Zhang, Yihan Wan, Yunhao Tang

cs.SD, cs.AI, eess.AS

2025-07-18

Mistral ships 3B and 24B audio chat models: Whisper encoder plus Mistral decoder, beating closed ASR on English short-form and Common Voice, Apache 2.0, runnable locally.

What problem this solves

Speech models split into two camps. Whisper-class systems transcribe well and cannot answer questions. Closed audio chat models such as GPT-4o Audio and Gemini 2.5 Flash can do both, but the weights stay private and local deployment is hard. Missing is an open, chat-capable model whose transcription is strong and whose context covers tens of minutes of audio.

Benchmarks are skewed the same way. Most papers report WER and BLEU for transcription and translation, and almost nothing standardized for "listen, then QA / summarize / call a tool." Voxtral tries to fill both gaps: Mini and Small under Apache 2.0, plus speech-synthesized GSM8K, TriviaQA, and MMLU.

Method

Three parts. The encoder is Whisper large-v3: 128 Mel bins, 50 Hz, a hard 30-second receptive field. Longer files are encoded in independent 30-second chunks with reset positions (chunked attention), then concatenated. Short clips are still padded to a multiple of 30 seconds; dropping padding cost 0.5 WER on FLEURS French, so padding stayed.

50 Hz is too long for the decoder: 30 minutes is 90k tokens. An MLP adapter downsamples 4× to 12.5 Hz. At that rate ASR barely moves and Llama QA is about 1.5 points above the 50 Hz baseline; 6.25 Hz costs more than a point of French WER. Mini sits on Ministral 3B (4.7B total, 640M encoder). Small sits on Mistral Small 3.1 24B (24.3B total). A 32K window is about 40 minutes of audio.

Pre-training mixes two templates equally. Repeat: audio followed by its own transcript. Continue: audio followed by the next text segment, interleaved, for cross-modal discourse. <repeat> and <next> disambiguate. Repeat-only kills Llama QA. Continue-only sends ASR WER near 60%. A warmup that trains only the adapter helps understanding. Post-training synthesizes QA, summaries, and translations from long-form transcripts with Mistral Large, and TTS-converts text SFT data. Pure TTS overfits synthetic speech, so they also cut real ASR questions that world knowledge can answer. Transcription gets its own special token. Alignment uses DPO and online DPO: sample two replies, swap audio for a transcript, score with a text reward model.

Results

Transcription is the hard number. Voxtral Small beats Whisper large-v3, GPT-4o mini Transcribe, Gemini 2.5 Flash, and ElevenLabs Scribe on macro English short-form and Mozilla Common Voice. LibriSpeech test-clean: 1.53 versus Whisper 1.84. SPGISpeech: 1.89 versus 3.15. On 10-minute Earnings-21/22 slices, Scribe still wins (7.39/9.16 versus 9.55/12.48). On every reported FLEURS translation pair Small leads; en→fr BLEU is 57.3 against 52.7 for GPT-4o mini Audio.

On understanding, Small trades blows with closed models in the same price class: Llama QA 71.7 vs 74.3, Openbook QA 88.4 vs 83.7, spoken GSM8K 89.7 vs 90.8. On the internal SU benchmark, Small SFT already scores 86.61% helpful / 4.16 grade; online DPO reaches 88.31 / 4.38 but English short-form WER rises from 6.31 to 6.50, so the public Small checkpoint stays SFT. Mini ships the online-DPO weights. Text scores stay close to Mistral Small 3.1.

ModelLS-clean WERSPGI WERSU helpful
Whisper large-v31.843.15
GPT-4o mini Transcribe1.924.51no transcribe mode
Gemini 2.5 Flash2.974.0088.64%
Voxtral Small1.531.8986.61% (SFT)

Why it matters

Few open audio models are local-runnable, permissively licensed, and unified for chat plus transcription. 32K context turns "upload a 40-minute meeting and ask" into a model feature. Audio-native function calling shortens the usual cascade for voice agents. The synthesized MMLU/GSM8K/TriviaQA sets are the missing evals, and they released them.

This is still a cascade in spirit: Whisper listens, Mistral thinks, text comes out. It shows the 24B cascade can beat same-tier closed ASR. It does not do full duplex or paralinguistics.

Limitations

Arabic Common Voice WER stays above 45% for every model, about 62% for Small, so Arabic is dropped from the macro-average. The SU judge and the DPO reward model both see transcripts, not waveforms, so emotion, accent, and ambient sound are out of scope by construction. Online DPO helps answers and slightly hurts English short-form WER; Small therefore does not ship those weights. Padding, chunked encoding, and synthetic QA are engineering compromises with no ablation against a newer encoder. Comparisons with GPT-4o and Gemini are for one price band and one eval suite.

Terms

Source

What people are saying

Related papers

All paper explainers