MiniMax's Large Model: 3 Billion Daily Interactions with the World | Oasis Vitality

I need more context to translate this accurately. "参赞生命力" could mean several things depending on context: - **"参赞" as a diplomatic title**: "Counselor [on] vitality/life force" — but this would be an unusual pairing - **"参赞" as "participate in and support"**: "Engaging with and nurturing vitality" or "Contributing to life force" - In a **wellness/health product or brand context**: could be something like "Vital Counsel" or "Vitality Advisor" Could you provide the surrounding text or context? For example: - Is this a product name, article headline, or section title? - What topic is the full article about (health supplements, traditional Chinese medicine, a diplomatic event, etc.)? - What comes before and after this phrase?

August 31 marked MiniMax's first-ever "Partner Day." At this offline launch event, MiniMax founder Junjie Yan openly reflected on the company's entrepreneurial origins, while sharing a key user metric — 300 million daily AI interactions, processing 3 trillion tokens. Oasis Capital, one of MiniMax's Series A investors, is thrilled to see the company's achievements to date. Jinjian Zhang of Oasis Capital also sent his well-wishes in an opening video: "The first time I met Junjie, we sat in a teahouse chatting for three or four hours. He told me about many opportunities he saw. What MiniMax has always emphasized is 'user in the loop.' A large language model isn't an atomic bomb, nor is it a rocket. It's not some highly esoteric technology. Ultimately, it's a product that serves tens of millions of users."

Looking back, Yan's early emphasis on "user in the loop" and the now-frequently mentioned "Intelligence with Everyone" are cut from the same cloth — his unwavering belief has always been that the ultimate goal of AI development is to become more general-purpose and help everyone. Below is a recap of MiniMax's first Partner Day. Enjoy.

Two and a half years after its founding, MiniMax is speaking to the public for the first time about its original aspirations and vision. MiniMax's large models now handle 300 million daily AI interactions with users worldwide, processing 3 trillion tokens per day. The Partner Day also unveiled applications of a new-generation model technology based on MOE (Mixture of Experts) + Linear Attention, alongside multiple latest R&D achievements. MiniMax answers "300 million daily interactions with the world" through technological breakthroughs, with "Intelligence with Everyone" as its vision, believing this to be the most efficient — and only — path to AGI. At this Partner Day, MiniMax released video model abab-video-1, music model abab-music-1, and updated speech model abab-speech-1. Additionally, MiniMax will release abab 7, a multimodal model using MoE + Linear Attention technology, within the coming weeks. MiniMax founder Junjie Yan delivered a keynote sharing more highlights, presented below.

Latest Technology Releases

Video model abab-video-1 abab-video-1 features high compression rates, strong text responsiveness, diverse styles, and support for native high-resolution, high-frame-rate video with cinematic quality. Users can log into the Hailuo AI web version (www.hailuoai.com/video) to try generating their own exclusive creative videos. Music model abab-music-1 abab-music-1 supports multifunctional end-to-end music generation, capable of synthesizing various musical forms including instrumental music and a cappella works, with simultaneous accompaniment and vocal generation, greatly simplifying music recording and creation processes. Speech model abab-speech-1 The updated abab-speech-1 supports multiple languages including Cantonese, Korean, Spanish, and Japanese, with hyper-realistic generated speech and nuanced, natural emotional variation. New-generation MOE + Linear Attention model technology release Supports efficient training on massive data, with greatly improved practicality and response speed, significantly reducing training and inference costs for large models. Compared to general Transformer architectures, the new architecture reduces costs by over 90% at 128K sequence lengths, with advantages becoming more pronounced as sequence length increases. Multimodal model abab 7 using MoE + Linear Attention technology will launch within weeks. In capability comparisons with same-generation models like GPT-4o, the new-generation abab model doubles efficiency when processing 100,000 tokens, with greater improvements as length increases.

MiniMax is China's large model company with the highest daily processing volume and interaction duration

MiniMax's large models conduct 300 million daily interactions with global users, including:

  • Over 3 trillion text tokens processed daily, equivalent to experiencing 3,000 lifetimes in a single day;

  • 20 million images generated daily, equivalent to the painting collections of 400 Forbidden Cities;

  • 70,000 hours of speech synthesized daily, equivalent to reading 7,000 books in one day.

Intelligence with Everyone

Often, it's not our technology helping users — it's users helping us.

Our users are diverse, creative, and vibrant. It's through their participation and feedback that better intelligence emerges. Technological innovation and breakthroughs stem from people's faith in and commitment to technology. MiniMax will continue driving accelerated AI development through technological breakthroughs, expanding the boundaries of human intelligence. Currently, MiniMax's video and music models are live on both the Open Platform and Hailuo AI web version. We hope to collaborate with more ecosystem partners to explore落地场景 for music and video models, empowering creativity and imagination through technology. Working together with all of you, we strive to truly achieve Intelligence with Everyone.

Full transcript of IO's speech:

Hello everyone, I'm IO, founder of MiniMax. Welcome to our first Partner Day event.

First, let me share the story before MiniMax's founding. Before starting this company, I spent over a decade on AI R&D. What was AI back then? The most representative applications were facial recognition and AlphaGo. In the past, most scenarios required custom models, yet you couldn't customize for every scenario. So for many people, AI remained just a lofty concept. This increasingly puzzled me as a practitioner: we put so much effort into AI research, but what was it really for? During the 2021 Spring Festival, I returned to my hometown to visit my grandfather. Their generation's life experiences were my favorite stories growing up. My 80-year-old grandfather wanted to write a memoir, but he couldn't type and lacked the energy to research materials. Theoretically AI should have been perfect for this, but unfortunately, AI at that time couldn't do it.

This made me realize that the ultimate goal of AI development is to become more general-purpose and help everyone. Three words summarize this: Intelligence with Everyone. Once I understood this, everything became clear. It restored my original passion for AI research and a strong sense of mission. But questions followed: how to begin?

To pursue this goal, we founded MiniMax at the end of 2021. In a room smaller than 100 square meters, we wrote down our original aspirations and path. Three judgments from that day remain firm choices to this day.

First, we believe next-generation AI is an intelligence infinitely close to passing the Turing test — naturally interactive, within reach, everywhere.

Second, this is extraordinarily difficult. It looks hard now; it was even harder then. Achieving this goal is more like building chips — a massive systems engineering project. You can't settle for 5% or 10% improvements; you need technological breakthroughs that bring order-of-magnitude gains.

Third, because this is so difficult, we must firmly take it step by step, breaking down the problem. We judged we should start with high-tolerance-for-error applications like casual chat and writing. As technology improves step by step, we can build more powerful, problem-solving-oriented applications, ultimately bringing extended intelligence to everyone.

Intelligence with Everyone — co-creating intelligence with users — is not just the goal, but also the most efficient, perhaps the only, path. Often it's not our technology helping users; it's users helping us. With diverse users' participation and feedback comes better intelligence.

From our founding on December 9, 2021 to today, exactly 996 days have passed. Currently, MiniMax's large models and end users (including our own products plus Open Platform partners) conduct 300 million interactions daily. What's 300 million? It includes processing over 3 trillion text tokens daily, generating 20 million images daily, and generating 70,000 hours of speech daily.

What are 3 trillion text tokens? Equivalent to experiencing 3,000 lifetimes in a single day.

Behind these 300 million connections are users from around the world, users who've grown with us. They come from all across China, men and women, young and old, all sharing common traits — diverse, creative, and vibrant. We strive to use good technology to co-create surprise moments with them, which is our deeper motivation for focusing on improving technology. These users' real stories converge into over 300 million minutes of daily interaction time with MiniMax models.

Interaction duration is also the best proxy for processing volume, because on many third-party data sites like QuestMobile and Sensor Tower, you can find related data. A year ago today, our daily interaction duration was roughly 3% of ChatGPT's; today it exceeds 50%. This is also currently the largest interaction duration among all Chinese companies.

Multiple data points suggest we may be the Chinese large model company with the largest daily processing volume. But even with this progress, the users we connect with haven't reached 1% of the global population. There's still a long road to our Intelligence with Everyone goal.

How to grow from today's 1% to 100%? The most important thing is increasing AI product penetration and depth of use among users. Based on over two years of repeated reviews and summaries, we believe improving these two metrics can only be accomplished through one thing: "Science and technology are the primary productive forces." Looking at the narrow domain of large models — whenever our models make significant improvements, when processing speed notably increases, we can see user scenarios and depth of use significantly increase. Conversely, here's a real case that actually happened: a bug caused conversation repetition error rates to rise, and daily conversation volume dropped 40% that day. This also explains our most fundamental reason for persisting in technological innovation.

Today's AI applications still have many technical hurdles to overcome to achieve qualitative improvements in penetration and depth of use. We believe the three most important optimization directions are:

  1. How to continuously reduce model error rates: Current models still have relatively high error rates, sometimes stunning, sometimes unreliable. This also constrains models from handling complex tasks, because complex tasks often require multiple steps, and higher error rates lead to exponentially increasing failure rates. Reducing model error rates is the most fundamental prerequisite for enabling models to handle complex tasks, and also the core means of increasing user depth of use.

  2. Unlimited-length input and output: Why does this matter? The simple reason is that humans have this capability, able to process unlimited-length input and output. Traditional large models' computational demands rise quadratically with input-output volume, quickly reaching limits that computing power cannot sustain, requiring底层创新 to solve.

  3. Multimodality: From life experience it's not hard to see that text interaction is just a small part; more is voice and video interaction. Multimodal content like sound, images, and video has become the mainstream of information transmission. To increase penetration, multimodality is the necessary path.

So how to overcome these technical hurdles? In the large model field, we believe that within the same capability range, "fast is good." We all know about Scaling Laws in large language models, meaning that with the same algorithm, more training data and parameters mean better results. Therefore, between two similarly performing models, the one that trains and infers faster can more effectively utilize computing resources to iterate on more data, thereby achieving better model capabilities.

So we believe fast is good — a simple philosophy that's easily overlooked. "Fast" is MiniMax's core R&D goal for its underlying large models. Around this, we've made many technological innovations; here are two specific examples.

First, MOE. Before MOE architecture was industry-recognized, we made a decision to be the first in China to achieve core breakthroughs in the MoE algorithm technology路线. We compared Dense models with non-native MOE and native MOE. In our previous-generation model 6.5s, our MOE model was 3-5x faster than the Dense model. This is also why the abab 6.5s model can process billions of interactions daily — a very core reason. Our abab 6.5s is fast enough, so it achieved wide deployment.

In solving MOE problems, we encountered many technical challenges, but after investing significant effort to ultimately solve them, this strengthened our confidence in self-developed technology and our courage to face complex technical challenges. This courage enabled us to solve an even harder technical challenge in recent months.

Second, Linear Attention. This doesn't just bring an order-of-magnitude improvement; it's also a key step toward solving unlimited-length input and output. Simply put, Linear Attention finds a right-multiplication approximation to the left-multiplication in Transformer computations, transforming the quadratic growth relationship between input length and computational complexity in traditional model architectures into a linear relationship.

Although someone proposed this idea back in 2019, no one had ever made it work at large scale. Our team found a new normalization method to replace Softmax, and a positional encoding to provide computational non-linearity. Beyond this, we found an efficient way to make large-scale training of Linear Attention possible.

Starting in April this year, as the first Chinese AI company to delve into Linear Attention, we successfully developed a new-generation model based on MOE + Linear Attention that can truly match GPT-4o levels. Taking three internationally leading models as examples — GPT-4o, Claude 3.5 Sonnet, and abab 7 — you can see that as input length increases, speed improvements compared to non-Linear Attention models change very significantly. When processing 100,000 tokens, the new model's processing efficiency reaches 2-3x improvement, and as length increases, model efficiency improvements become more pronounced. Theoretically, the model can process tokens approaching unlimited length.

In working on Linear Attention, we were pleasantly surprised to discover that GPT-4o actually does the same thing. This gave us great confidence — we found that in exploring frontier technology, we converge with the best international companies despite different paths. The MiniMax team has increasingly strong technological innovation capabilities. We need to persist, continuously finding innovations that accelerate technological progress, to truly have a chance at becoming a globally top-tier technology company.

We realize that even after achieving MOE, Linear Attention, and several-fold improvements, many other technological innovations still await us. Only with multiple technologies each bringing several-fold improvements, then multiplying them together, can AGI become reality. The abab 7 model's core technology is precisely based on MoE + Linear Attention. Beyond this, we've built multimodal understanding capabilities on abab 7. Additionally, we've applied similar innovative technologies across multiple models including text, voice, and video.

Today, MiniMax's speech model adds internationally leading and highly practical features:

  1. Multiple languages: Supports over 10 languages including Japanese, Korean, Spanish, French, and Cantonese. MiniMax has also become the world's first company with authentic Cantonese speech model capabilities;

  2. Emotional expression: Generated speech is hyper-realistic with nuanced emotional variation;

  3. Music: MiniMax's first music model has debuted. This model has extremely high artistry and plasticity, and we believe it will bring many new玩法 and surprises to our creators and partners.

Our speech model was honed through products like STARFIELD, Hailuo AI, and Talkie. We insist on using the same models in our own products and API.

MiniMax has launched our first video model, possibly the best video generation model currently available. Compared to video models on the market, our model's unique advantages are:

  1. Strong text responsiveness: Thanks to MiniMax's continuous accumulation in text, with good instruction following;

  2. High compression rate: Thanks to our accumulated experience in network architecture, with good expressiveness for highly dynamic, rapidly changing information, where Linear Attention's high inference efficiency contributes significantly;

  3. Diverse styles: We have globally diverse user distribution. Whether 3D cinematic blockbuster scenes or 2D animation, it's all驾驭able; whether Chinese style, sci-fi, or American comics, nothing stumps it.

When we combine updated, more powerful MiniMax model capabilities, what happens? We tried using multiple models to generate a short film The Magic Coin, with zero human modification. Going forward, we'll publish the prompts behind the video, providing a reference for "how to generate high-quality video content using only models." In the future, we'll同步 all newly launched models and capabilities on the MiniMax Open Platform and in STARFIELD and Hailuo AI for experience.

I have a magic coin here. We very much hope our AI can be like this magic coin — helping many people create infinite imagination, bringing it to everyone.

Regarding model and product updates, the speech model, music model, and video model are now fully released. Additionally, a new version of model abab 7 that can match GPT-4o in both speed and quality will be released in the coming weeks.

All models, including our best music model, speech model, best video model, and what we believe may become the best text model, can be experienced on the Open Platform. Our Open Platform currently has over 30,000 enterprise customers and developers, and continues to grow rapidly. Meanwhile, these models can also be first experienced in Hailuo AI, which is our flagship product in the personal assistant domain. When complex models are used together, what complex, more advanced玩法 can be created? We'll place this in the content community product STARFIELD APP.

As idealistic yet grounded MiniMax people, we continue striving forward. Two and a half years later, we're fortunate to have so many of you here today among our fellow travelers, along with our growing global user base. Thank you all for your continued attention and support. We hope to work together with all of you, alongside MiniMax, to push the boundaries of human intelligence a little further, truly achieving Intelligence with Everyone.