Mistral Launches Open-Source Speech Model for Enterprise Voice AI
French AI innovator Mistral has unveiled a new open-source text-to-speech model designed for both voice AI assistants and enterprise applications, including customer support. The release positions the company to compete directly with AI players like ElevenLabs, Deepgram, and OpenAI.
The model, named Voxtral TTS, supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic, enabling global deployment for multilingual voice applications.
Compact, Cost-Effective, and High Performance
“Our customers have been asking for a speech model. So we built a small-sized speech model that can fit on a smartwatch, a smartphone, a laptop, or other edge devices. The cost of it is a fraction of anything else on the market, but it offers state-of-the-art performance,” explained Pierre Stock, VP of science operations at Mistral AI, in a phone interview with TechCrunch.
Voxtral TTS can create a custom voice using less than five seconds of audio and captures subtle speech nuances, including accents, inflections, intonations, and irregularities. Built on the Ministral 3B model, it can switch seamlessly between languages without losing the original voice characteristics—a feature particularly useful for dubbing or real-time translation. Stock emphasized that the goal was to make the AI sound human, not robotic.
Real-Time Performance Metrics
Mistral designed Voxtral TTS for real-time applications. The model achieves a time-to-first-audio (TTFA) of 90 milliseconds for a 10-second, 500-character sample. Additionally, with a real-time factor (RTF) of 6x, a 10-second audio clip can be generated in roughly 1.6 seconds.
Earlier this year, Mistral introduced two transcription models: one optimized for large batch processing and another for low-latency real-time use. With Voxtral TTS, the company appears to be building a comprehensive suite of enterprise voice solutions.
Towards an End-to-End Multimodal Platform
“We plan to have an end-to-end platform that can handle multimodal streams of input, including audio, text, and image and output as well. The main benefit of that is you get way more information with an end-to-end agentic system that supports audio as an input or output,” Stock added.
Mistral believes that its open-source approach and customizable features will make it attractive for enterprises, allowing businesses to tailor voice models to their specific needs and differentiate from competitors.





0 Comments