MiniMax Voice Model Claims Global Top Spot | Oasis Capital on Vitality

Counselor on Vitality

AI voice is approaching its Her moment.

As countless agents and hardware devices enter our lives in unprecedented ways, AI voice interaction is seeing explosive growth. The personalized demands of massive numbers of endpoints, customers, and creators all need to be met at scale by the same underlying model. Beyond natural, warm voice experiences, "personalized voice" must be solved.

Current leading text-to-speech (TTS) models, while impressive, typically offer only limited voice and language options. This not only restricts user choice but fails to capture the cultural diversity embedded in human language.

We've developed a high-quality TTS system based on the AR Transformer model — MiniMax Speech 02. The model has strong enough generalization capabilities to easily handle 32 languages, different accents, and different emotional tones.

The core innovation of this model system lies in its intrinsic Zero-Shot capability, which we've named Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder. Architecturally, we've designed a "learnable voice extractor" that can flexibly collaborate with the AR Transformer.

We trained it jointly with the AR Transformer, significantly improving speech synthesis quality. Because of this, we can offer infinite combinations of any language × any accent × any voice through a single model, greatly enriching the diversity of voice generation. On the authoritative international platform Artificial Analysis, MiniMax Speech 02 also ranked first globally through user evaluations worldwide.

Dual Authority Rankings, First Globally

On two global authoritative speech benchmark leaderboards — Artificial Analysis Speech Arena and Hugging Face TTS Arena — MiniMax Speech (listed as Speech-02-HD) surpassed top-performing global models including OpenAI and ElevenLabs, ranking first on both. Beyond professional metric evaluations, the Arena leaderboard's ELO scores are derived from users randomly listening to and comparing voice samples from different models, then selecting the better one; the results prove that in terms of user experience, MiniMax Speech 02 delivers superior listening quality.

Artificial Analysis Speech Arena Leaderboard

Hugging Face TTS Arena Leaderboard

While delivering superior listening quality, MiniMax Speech 02 also achieves lower pricing — half that of ElevenLabs Flash V2.5 and one-quarter that of Multilingual V2.

Flexibility from Model Architecture

The "learnable voice extractor" is essentially a speaker encoder, capable of converting audio clips of arbitrary length into fixed-size conditional vectors, thereby achieving high-quality, flexible voice expression.

  • Zero-Shot for hyper-realistic voices: Requires only a reference audio clip, with no corresponding text needed; in this Zero-Shot approach, the encoder extracts voice characteristics solely from the reference audio, thus better capturing the essence of sound — timbre, pitch, and style — enabling more flexible and extensive decoding space for prosody. The final output rivals real human speech while being more stable than actual humans.
  • High-quality synthesis in 32 languages: During reference audio processing, the speaker encoder handles voice characteristics decoupled from semantic content; because the speaker encoder is learnable, it can be trained on all languages covered in the training dataset. This is why MiniMax Speech fundamentally supports 32 multilingual languages with superior cross-lingual performance.
  • Scalable features and personalized expression: Since the conditional vectors achieved by the speaker encoder are themselves decoupled, this gives MiniMax Speech downstream application scalability flexibility. We've implemented flexible emotional expression in arbitrary voices, voice generation from voice descriptions, and clone enhancement based on specific speakers. These features further enrich MiniMax Speech's personalized voice space.

For more technical details, experimental comparison data, and the open-source multilingual test set, please read the technical report.

GitHub:

https://github.com/MiniMax-AI/MiniMax-AI.github.io/blob/main/tts_tech_report/MiniMax_Speech.pdf

Hugging Face:

https://huggingface.co/spaces/MiniMaxAI/MiniMax-Speech-Tech-Report

Showcase

Charismatic presentation voice

"What if I told you the best performing marketing strategies right now are the exact ones most experts would warn you to not even try? You've been told to follow the rules, play it safe, stick with what's proven. But here's the twist. Some of the weirdest, most backward-sounding marketing tactics out there are quietly crushing everything else. And yes, there's data to back it up……Pretty much every marketing agency out there has tested this over and over and time and time again. Ugly or amateur-looking Facebook and Instagram ads often get significantly better click-through rates and lower cost per click. And it's not just a fluke, it's a pattern."

Intersubjectivity theory in ASMR style

"Habermas's theory of intersubjectivity, its core is the paradigm of communicative rationality~ He believes that truth must be achieved through effective dialogue between subjects to reach consensus~ Oh~ The ideal speech situation needs to be built upon four validity claims: comprehensibility, truth, rightness, and sincerity."

Multilingual

Thai: "สวัสดีค่ะ วันนี้อากาศดีมากเลย คุณจะไปทานอาหารกลางวันที่ไหนคะ ฉันกำลังคิดว่าจะไปร้านอาหารไทยแถวนี้"

Polish: "Młoda sowa siedzi cicho na gałęzi sosny, obserwując leśną polanę w świetle księżyca. Wiatr delikatnie porusza liśćmi drzew."

Japanese: "電車が遅延している影響で、渋谷駅がとても混雑しています。次の山手線は約10分後に到着予定です。お急ぎのお客様は、他の路線もご利用ください。"

Zero-Shot cross-lingual output case

Japanese + Korean: "最近の天気予報によりますと、今週末は桜の開花に最適な気温になる予定です。東京都内の各公園では花見客で賑わうことが予想されますが、서울에서도 벚꽃이 피기 시작했다고 하네요. 이번 주말에는 여의도 공원에서 벚꽃 축제가 열린다고 하니 많은 분들이 찾아오실 것 같습니다."

English + Chinese: "Kiddo! Come come come, 学如逆水行舟,不进则退。I see you're using AI tools already - so smart! But eh, cannot just rely on tools only lah! The future belongs to those who can work alongside AI, not those scared of it."

English + Spanish: "Mi abuelita always told me "el que persevera, alcanza". If you persevere, you'll achieve your dreams!Guess what! They choose me to play the lead role in our BIG show!"

Voice from text description

Voice description: English-speaking middle-aged male voice, slightly husky, speaking at a moderate-to-slow pace with a deep tone. Like someone telling an old story, conveying a nostalgic feeling, with a relaxed and composed manner of speaking.

"That was back in the late 1970s. I remember when our village first got electricity - everyone was so excited. In the evenings, people would bring their stools and gather under the big banyan tree by the village committee office to watch movies projected on the wall. Even now, thinking back to those moments still fills me with warmth."

Voice description: Voice of a young Chinese woman, crisp timbre, relatively fast speaking speed, lively intonation, like doing a game livestream, voice carrying a cheerful feeling with overall higher pitch, overall relaxed atmosphere.

"啊!这里有个宝箱!让我们看看里面是什么~哇!是传说中的紫色装备!运气也太好了吧!谢谢小伙伴们的打赏,我们继续往前探索......"

Visit the MiniMax Audio page to experience the powerful capabilities of MiniMax Speech:

https://www.minimax.io/audio

https://www.minimaxi.com/audio

Multilingual Benchmark MiniMax Speech supports synthesis in 32 languages. To evaluate its multilingual performance, we built a dedicated test set and conducted comparative evaluation against ElevenLabs's multilingual_V2.

  • Both models clone voices and generate in Zero-Shot fashion;
  • For WER (Word Error Rate) calculation, Whisper-large-v3 or paraformer-zm was used for transcription;
  • SIM (voice similarity) is determined by calculating cosine similarity between speaker embeddings.

Test results show:

  • On the SIM (voice similarity) metric, MiniMax Speech 02 outperforms ElevenLabs across all languages; this indicates that MiniMax Speech 02 has superior multilingual expressiveness under Zero-Shot conditions.
  • MiniMax Speech 02 demonstrates excellent accuracy in mainstream European and American languages including English, French, Italian, and Portuguese. In contrast, ElevenLabs's word error rate exceeds 10% on some Asian languages such as Cantonese, Thai, Vietnamese, and Japanese. This fully demonstrates that MiniMax Speech is more powerful and reliable in multilingual adaptation.

Enhancing Voice Quality To optimize the quality of generated speech, we use Flow-VAE to compress audio into latent features, and model these latent features through a Flow Matching model.

Traditional VAE typically assumes a standard normal distribution for the latent space, while Flow-VAE introduces a flow model. This approach constrains the encoder output distribution to a normal distribution rather than a standard normal distribution, thereby enhancing the encoder's information expression capability.

Flow-VAE provides richer audio representations than traditional mel spectrograms; Flow Matching can accurately model the distribution of these audio representations. Combined, they enable MiniMax Speech 02 to express more details in generated speech. In listening experience, this brings higher audio quality and higher similarity.

Going forward, we will continue to improve the model's controllability and efficiency.

Overseas, we already support numerous content creators who use low-threshold voice tools to flexibly take orders with their own voices, performing voice acting for advertisements and short films, empowering the gig economy. Additionally, through support for scarce and precious minority languages, MiniMax hopes to use AI to transmit multilingual voices to the world with the most authentic local pronunciation, so that every language globally is heard and every culture is understood.