A Conversation with Junjie Yan of MiniMax: AGI Isn't a Doomsday Weapon, It's a Product Ordinary People Use Every Day | Oasis Vitality

Counselor Vitality

In 2022, before anyone was paying attention to GPT, MiniMax founder Junjie Yan saw the boundless potential in what was coming and resolutely set out to build in that direction. Later, when everyone found themselves in the spotlight of large models, eager to speak out and make noise, he was quietly exploring what products could actually work. That steadiness and conviction — the ability to stay grounded and believe — is exactly the founder quality that most moved Oasis Capital. Below is Junjie Yan's first-ever public interview. Enjoy. This article is republished with permission from LatePost (ID: postlate). By Manqi Cheng. Edited by Song Wei.

In China's bustling large-model startup scene, MiniMax founder Junjie Yan may be the most mysterious figure of all. He has never appeared in public, never given a single interview. Even with his company valued at $2.5 billion, he remains silent.

Among high school and college students, products like Glow and STARFIELD have become wildly popular — they're playing romance simulation games, chatting with virtual beings, writing unhinged rants, even building cities. But few realize they're all powered by MiniMax.

Following its recent Series B, MiniMax has become one of China's highest-valued large-model companies, with both Tencent and Alibaba on its shareholder roster. It has only been around for two years.

"Back then ChatGPT hadn't launched, OpenAI was still lukewarm, no other Chinese companies had emerged yet, and SenseTime was about to go public — Junjie Yan jumped out to start a company. He truly has faith in AGI," said Mingming Huang, founder of Future Capital and an early investor in MiniMax.

By then Yan had already served as Vice President at SenseTime, Deputy Dean of its Research Institute, and CTO of its Smart City Business Group. Earlier still, he researched computer vision at the Chinese Academy of Sciences and Tsinghua University. Typically, technical talent shines in just one narrow domain, but Yan excelled across multiple fields.

Now 35, Yan is a soft-spoken manager who's always smiling. Investors describe him as an excellent listener, skilled at putting conversation partners at ease — only to quietly drop sharp observations once they've relaxed.

Junjie Yan presenting at the 2021 World Artificial Intelligence Conference Algorithm Competition Finals

In founding MiniMax, Yan showed none of the hesitation typical of first-time entrepreneurs. He seemed willing to make big bets.

In the second half of 2023, while most Chinese peers continued iterating on dense models — the more reliable path to incremental performance gains — Yan poured nearly all of MiniMax's R&D and compute resources into something far more uncertain: MoE (Mixture of Experts) models. His judgment was that if they were going to serve tens of millions or even hundreds of millions of users, they had to go MoE. Otherwise, "the cost and latency of generating tokens would be unacceptable — we'd collapse quickly."

Before the large-model frenzy began, MiniMax had secured substantial GPU compute from ByteDance's Volcano Engine at relatively favorable rates, stockpiling ammunition. Yet MiniMax doesn't own any GPUs itself. Yan believes holding assets would only distort decision-making.

"The technical path I chose has the highest ceiling, with almost no retreat," Yan said. "The compute approach I chose is equally aggressive."

Robin Li believes the "dual-wheel drive" of simultaneously building models and applications isn't a good model for startups. But from day one, Yan believed that for a large-model startup to grow independently and become significant, both technology and product had to be excellent.

MiniMax was the earliest, most prolific, and most heavily invested among China's large-model startups on the product front: of roughly 300 people, nearly 200 work on product. Their first product, Glow, launched in October 2022, followed by STARFIELD, Hailuo AI, and at least four others — spanning companion-style social entertainment apps to productivity tools like Q&A. Multiple apps have surpassed one million DAU.

Yan believes startups have only one path to independent growth: build massive 2C products before the window of rapid technological evolution closes.

"Without product to capture it, even if you make a technical breakthrough, it ultimately won't be yours," Yan said.

China's market now has six large-model unicorns (MiniMax, Moonshot AI, Zhipu AI, Baichuan AI, 01.AI, StepFun), with more than 100 models in total. They're racing to catch GPT-4 — released a year ago — with fewer resources, while navigating fierce competition.

Some prioritize survival first. Others believe "Go big or go home," refusing to exist in the middle. Yan is the latter.

Below is LatePost's conversation with MiniMax founder Junjie Yan.

"Everything only gets good when taken to the extreme"

LatePost: An OpenAI engineer told us his test for whether an AI founder has genuine AGI conviction is whether they started their company before or after ChatGPT launched.

Yan: MiniMax was founded at the end of 2021. At that time, AGI was still a massive non-consensus in China.

We calculated back then that scaling GPT-3 by 100x would require an enormous amount of money — possibly tens of billions of dollars. But at that point, we certainly didn't think China would have that much capital willing to back a startup.

LatePost: Some people think you started out doing metaverse and pivoted to AGI once large models became hot. How much did you actually believe in AGI at the outset?

Yan: We were founded before ChatGPT. Most companies came after. That's the core distinction.

Before ChatGPT, there were no reference points. You had to experiment more. But the innermost kernel was always technological progress — what was uncertain was product direction.

Our initial vision for AI products was an intelligent agent with simultaneous voice, visual, and text capabilities. We built a version with 3D avatars, somewhat like digital beings in the metaverse, but its language and speech capabilities were still powered by large models.

LatePost: What do you think AGI actually is? If one day it's realized, how would we know it has arrived?

Yan: We had a fuzzy definition back then, and it barely changed: the day people stop thinking of AI as AI, that's probably when it's here.

Just like when we talk about Douyin today — you don't think of it as a content distribution software based on recommendation algorithms. Douyin is just Douyin.

LatePost: MiniMax was the first Chinese company to say it would do AI 2C. Why?

Yan: Before deciding to start a company, I kept thinking about what kind of technological progress could generate sufficiently high feedback for society. What came to mind were electric vehicles, mobile internet. The biggest characteristic of these industries was serving ordinary people. And the prerequisite for serving ordinary people was commercialization — it had to be a product, not a project.

At the time, the entire AI industry was stuck in a rut, while the industries that had truly succeeded operated completely differently. The conclusion was almost singular — we needed to build sufficiently productized AI technology and products that could serve the masses, not projects serving a handful of big clients.

So I never believed AGI would be like an atomic bomb, some doomsday weapon. It's something ordinary people use every day — a product, a service. That's what we've held to most firmly.

And AGI shouldn't be built by one company alone. It requires that company and its users to build it together.

LatePost: In January this year, you were the first in China to launch an MoE large model. Other companies were mainly iterating dense models last year because progress was faster and more certain. Was betting on MoE a gamble?

Yan: At first I did think we were gambling. Those months, others were advancing rapidly on steadier ground, while we were betting on something harder.

We allocated over 80% of our compute and R&D resources to MoE, with no Plan B.

LatePost: MoE development started in summer 2023. Why was it absolutely necessary then?

Yan: First, we knew exactly what baseline resources and data we had. Based on that compute and data, only MoE could actually finish training — from the ceiling of what you could train, it had to be MoE.

Second, we already had many users by then, with both B2B and B2C products. Our models were processing massive amounts of tokens daily. We realized that if we kept doing dense models, the cost and latency of generating tokens would be unacceptable — we'd collapse quickly. So MoE was the only option.

Of course this is probably industry consensus now: if you're building a trillion-parameter model, you can't do it as a dense model.

LatePost: How did you finally crack it?

Yan: The process was painful. We failed twice. We already had plenty of uncertainty, and doing something new added more — challenges were inevitable.

For instance, you'd train a model for half a month and find certain metrics drifting further from initial estimates. Like launching a rocket expecting 30,000 meters, but it veers off course. You start diagnosing what went wrong, fix the problem, and find you're still not in a good state — another failure. But you accumulate experience, synthesize it, and try again.

Each attempt cost a lot of money. More importantly, time.

I later realized it wasn't really gambling. Many challenges weren't inherent to MoE itself, but stemmed from more fundamental issues: experimental methodology, network and data structure exploration, and so on.

Ultimately, solving problems didn't come from "solving MoE" per se, but from identifying past shortcomings that made the entire R&D team more efficient and rigorous.

LatePost: Someone who's worked with you described you as having strong engineering thinking — you pursue optimal outcomes under given constraints.

Yan: It's all calculated. Most decisions at our company are based on optimizing certain variables — we're solving equations.

LatePost: These days, every company's resources — your constraints — are shifting rapidly. When you run the numbers, do you tend toward conservatism or risk?

Yan: We basically always choose the most aggressive option, because you only get good results when you take things to the extreme.

The technical path I chose also has the highest ceiling, with almost no fallback. Our approach to compute is similarly aggressive.

LatePost: I heard you don't buy GPUs, you only rent.

Yan: We don't own a single GPU, even though we're probably the Chinese startup that actually uses the most GPUs.

Because owning assets distorts your behavior. If I had a lot of GPUs, the commercially optimal thing would be to rent them out to others. I want to keep the company simpler than that.

LatePost: Last October you ran into a compute crunch. How do you avoid similar risks?

Yan: Become the biggest customer in the market.


For Chinese startups, the better approach is to think about technology and product simultaneously

LatePost: Robin Li said the "dual-wheel drive" model isn't good for startups, yet you decided to build products from day one. How did you make that call?

Yan: When you first start out, you don't really get to think about such things — you have no technology, no product, no users. The first six or seven months were just about building the most primitive model. Only then did products follow.

If everything were free, if you had an infinitely strong organization, then having great technology would matter most, because you'd already have users, traffic, and monetization capabilities — you could rapidly test many products.

But that's not the case for startups. Without strong enough product capabilities to capture value, even technical breakthroughs won't ultimately belong to you. An independently developing startup must think about product.

LatePost: OpenAI only started building the killer app ChatGPT after GPT-3.5. Before that, OpenAI didn't prioritize product as much.

Yan: That's because OpenAI had order-of-magnitude leads in technology, talent, and data accumulation, which gave it a roughly one-year startup window. I don't think any company in the world will get such a unique window again.

No one will be 10x OpenAI. No one can suddenly produce something ten times better than everything else worldwide.

From this it follows that for startups — at least for Chinese startups — the better approach is to think about technology and product simultaneously.

LatePost: Some investors think you're building products too early. "You can't build Douyin on a BlackBerry."

Yan: By that logic, you shouldn't work on technology now either, since current technology won't be the technology of five years from now.

But clearly everyone agrees we need to work on technology now: only by building today's technology can you deeply understand it, and only then can you build the technology of three or five years from now.

LatePost: Technology develops incrementally. Is product the same? Products of this era are completely different from the last era's.

Yan: Product too. Many successful Chinese companies — miHoYo, Meituan, ByteDance, Li Auto — all share one trait: none of them succeeded with their first product. They all succeeded with their second or later products.

This wasn't my observation; a friend of mine summarized it.

LatePost: So why not just focus entirely on product? There are so many open-source large models now.

Yan: The core reason is that understanding the model is essentially equivalent to understanding the product. The deeper you go into product, the deeper your model understanding must become.

Another objective reason is cost and response time. Without strong control over the model, it's hard to control product cost changes or tune response times for users. And when building products, you encounter many problems: which ones are solvable? Which aren't? How do you iterate? These all require technical mastery.

One reality: last year many products were built on GPT-4. Why did no one create an experience rivaling ChatGPT?

LatePost: You're building products like others, but while some focus on one, you're building many simultaneously — Glow, STARFIELD, Hailuo AI, and others. Why a product portfolio instead of focusing on one or two?

Yan: OpenAI's products after ChatGPT weren't that successful either. If even OpenAI fails at product, it shows there's a gap between what product understands about technology and what technology can actually deliver.

The core issue is that even with the best technology and the best product, they won't match.

If you accept this gap, the objective law is: you should try more, fail more, and find what can actually succeed.

LatePost: Sounds a bit like how ByteDance builds products.

Yan: We're not qualified to operate like ByteDance yet.

Every company chooses the form most suitable for itself. For ByteDance, the most important thing is technical resources, because all its products are ready and it has unlimited product resources, so more attempts benefit it more. And each investment, even if the product fails, brings more experience and knowledge — the improvement for them is enormous.

We're the same. And compared to model R&D investment, product investment's resource share isn't that large. Based on our company's current state, you can calculate that this maximizes success probability.

LatePost: Technology matters, product matters — have you agonized over which matters more?

Yan: We agonized before, but not anymore.

In late 2022 when we were building Glow, we had a very painful experience. The team had all caught COVID, and the final release at the end of 2022 had a bug that degraded user conversation experience by about 15%. Our DAU dropped 40% over the three-day New Year's holiday. Finally, on the last day of the holiday, we couldn't take it anymore and found the bug — it was actually just one very small line of algorithm. We fixed it, and user numbers quickly recovered.

The lesson for us was that at this stage, product value fundamentally comes from your model performance and algorithmic capability.

We've gone through this several times. You can build many product features, but you'll find that almost all major improvements come from model progress itself.

LatePost: Building large models and so many products simultaneously — what's the biggest challenge?

Yan: The technology isn't good enough — that's the most fundamental issue. Our technology iteration speed is already quite fast, but we still have gaps versus the world's top models.


Tenfold Scaling Laws

LatePost: European AI leader Mistral has already open-sourced an MoE model, and the industry widely believes OpenAI's GPT-4 is also MoE. Will MoE be a battleground in large models this year?

Yan: MoE is just one piece; there are many other pieces. If something can be written in a single paper, you can basically assume it's not an absolute moat.

LatePost: In this technology race, what non-consensus judgments does MiniMax hold?

Yan: If there's any non-consensus in this industry, within 6-9 months it quickly becomes consensus.

There are three things everyone can see now: first, Scaling Laws; second, the compute and capital needed to achieve the same model precision may drop several-fold each year, because algorithms and publicly available academic research are increasing, and many people will do free exploration; third, focusing on improving data quality currently yields greater returns.

So from these three points — Scaling Laws, cost decline for same-precision models, and the importance of data quality improvement — you can basically derive some decisions that we and other companies make. I think it's relatively straightforward.

LatePost: How do you understand Scaling Laws? What possibilities does it reveal to you?

Yan: Scaling Laws is just a curve. You can believe in the original Scaling Laws, or you can believe in ten-times-faster, even hundred-times-faster Scaling Laws.

The 2020 paper that originally proposed Scaling Laws for large models, "Scaling Laws for Neural Language Models," argued that the most important variables affecting model performance were compute, data volume, and parameters, and gave the numerical relationship between these variables: C≈6ND, where C is Compute, D is Dataset, and N is Parameters. Model architecture and layer count had less impact on performance. It was more about providing a methodology: you could predict results of larger experiments through smaller-scale experiments. Second, it helped align the industry, because this endeavor requires division of labor across data, compute, chips, algorithms, and product — Scaling Laws gives everyone relatively consistent expectations. As for the formula and some conclusions in that paper, they may not hold up now. For example, it argued that layer count, architecture, and so on mattered less — at least several variables now appear to be important.

LatePost: Such as? What variables might let you achieve tenfold, hundredfold Scaling Laws?

Yan: For example, network architecture itself matters. When we worked on MoE, we initially thought good MoE architecture would resemble good dense architecture, but later found that wasn't the case — MoE itself can accelerate Scaling Laws. Also improving data quality. Also compute allocation — you can allocate compute to training or to data processing. Different choices could all accelerate Scaling Laws.

LatePost: Scaling Laws derives its power from its simplicity. When you introduce more variables, you break it.

Junjie Yan: There's no ceiling on improving data quality, optimizing algorithms, or refining training methods — keep doing them and they keep getting better. The real trade-off is that their efficiency gains for Scaling Laws accelerate at different rates in different cycles. But you can use small-scale experiments to predict which variables matter more at which stage — that's still the Scaling Laws methodology.

Why does China need to do several times the Scaling Laws? When compute is abundant, you can optimize the original Scaling Laws. When compute is scarce, you have to optimize a scaled-up version of Scaling Laws to achieve similar results.

This isn't impossible. Another Silicon Valley AI company, Anthropic, made Claude-3 comparable to GPT-4 in a shorter timeframe — that's essentially amplifying the original Scaling Laws. If one can do it, there will be a second, a third.

LatePost: Long Context is being discussed a lot lately. Will it become a differentiated path in the large model race?

Junjie Yan: A good large model should support long context by default. We've always had long context capabilities. We haven't emphasized this feature in our products mainly because of computational costs.

LatePost: What's the technical approach to achieving longer context processing?

Junjie Yan: Standard Transformers previously used non-linear attention. Over the past year or so, many people have been researching linear attention, which helps with long context.

The benefit of linear attention is that when text gets very long, its computational complexity grows linearly rather than quadratically. But actually, at 200,000 or 300,000 tokens, linear and non-linear perform similarly, because quadratic functions approximate linear functions early on. The difference only becomes pronounced at 800,000 to 1,000,000 tokens.

To my knowledge, Google's Gemini 1.5 was the first model to approach linear attention. If you call other APIs with very long text, the response is very slow. But Gemini 1.5 truly achieved something where 1,000,000 tokens compared to 500,000 tokens only takes 1x longer to respond, not 4x longer.

So long context doesn't solve problems at the 200,000 or 300,000 token level — it solves problems at 1,000,000 tokens and beyond.

LatePost: 1,000,000 tokens roughly equals 1,000,000 Chinese characters. How many people actually need this?

Junjie Yan: User demand and the capabilities you provide generate each other. Put a model here that far exceeds expectations, and gradually many people's needs will emerge.

For example, before ChatGPT had voice calls, no one would have said their need was voice calling. But once it was added, many people used it. Our voice conversation product — Hailuo AI's calling feature — is also quite popular. My grandfather is 80 years old. The first time he used this product, he discussed historical figures with it for 40 or 50 minutes. I never would have imagined someone using it this way.

LatePost: It seems you prioritized multimodal capabilities like voice in your products rather than long text. How do you judge which technical capabilities to optimize first?

Junjie Yan: We have a phrase: Intelligence with everyone. We're not the owner of this technology — that's our core belief.

Last year AI was very hot, but probably only 100 to 200 million people worldwide had used AI products, with only tens of millions being heavy users. Because asking good questions and following up continuously has a very high barrier. The people truly willing to type are probably just the people here in this room. More people are still accustomed to using voice.

We value multimodality because it allows more people to use AI, including elderly people and children. When we added images and voice to our products, we could clearly observe changes in user onboarding barriers and even penetration rates. The exact same thing already happened once in the mobile internet era, from Toutiao to Douyin.


The Further Along, the Higher the User Value

LatePost: Your first product, Glow, let users chat with customized AI characters, similar to otome games (romantic role-playing). It was very popular in the ACG circle. How did you come up with this direction?

Junjie Yan: When we did early product cold starts, we specifically targeted younger demographics — AI enthusiasts, ACG fans — and iterated the first few versions based on their experiences and feedback.

After we gained traction, we watched social media every day to see how users were using it. In our early product days, we didn't do A/B testing. We observed users, looked at user feedback, then verified with data and iterated.

LatePost: What pitfalls did you hit in building products?

Junjie Yan: Early on we worked on intelligent agents. Our vision was that they would simultaneously have voice, visual, and text capabilities. That's why the company built three models from the start — language, voice, and vision.

We quickly abandoned 3D avatars because they couldn't scale. The only major industries that previously used 3D were gaming and film, with development cycles of several years. At the same time, I realized using deep learning for 3D was the wrong approach.

On the current platform — mobile phones — if a 3D person is constantly looking at you, it's inherently strange. In most cases, interaction doesn't actually require a real visual form.

LatePost: Did you see this from certain data after launch?

Junjie Yan: Not data. When we made the first version of the avatar, we found two models to film. The moment we put 3D on a phone, we knew this was wrong.

LatePost: You hired a product manager before your first model was even built. How did you describe what kind of product you wanted?

Junjie Yan: I didn't know.

LatePost: You didn't know?

Junjie Yan: It was unclear at the time because there was no reference. We just imagined an intelligent agent that could have free, long conversations with you. Its essence was information exchange and processing.

What we could be certain of was that if the model was most important for serving the masses, it would definitely become a product. So we found a product manager very early on.

LatePost: Users have many demands. What do you satisfy and what don't you?

Junjie Yan: Our trade-offs became simpler over time: look at whether this demand aligns with technology development trends, and whether it can bring 10x or greater change to this type of user's experience.

LatePost: In terms of product taste, what do you think makes a good product? Your current products have many features — somewhat complex.

Junjie Yan: Honestly, we haven't made it yet, so there's no answer.

When you ask whether complex or simple products are better, most people will definitely say simple. But I'm somewhat skeptical of this, especially in the early stages of an industry. Consider that before Tencent made WeChat, they made QQ first, and QQ was a very complex product.

ChatGPT has about 30 million DAU and seems to have trouble growing beyond that. My conclusion is that a relatively simple AGI product, at the current technological stage, probably has an upper limit around there. But ultimately I believe there will be very simple interaction forms that satisfy broader needs.

LatePost: What inspiration did Sora (the text-to-video large model released by OpenAI) give you?

Junjie Yan: If Sora's response speed could become very fast in the future — generating a 1-minute video not in 20 minutes as it is now, but in real-time — that would be a huge change.

Then would it be a better video generation tool, or a better video generation community?

LatePost: A video generation community — the next step after that wouldn't be a super content platform?

Junjie Yan: Anything is possible. It depends on whether you believe this space is large enough, and whether you believe response times can become low enough.

LatePost: What do you think the AI product with the largest future user base might be?

Junjie Yan: We've only made products with millions of DAU. We haven't made tens of millions or billions-level products. Honestly, I don't know. I think it might still be information exchange and processing — its value is enormous.

LatePost: MiniMax's product DAU is already close to Character.AI (an American AI unicorn application where users can chat and interact with various AI characters), and time spent is even longer. But some people质疑 your good data isn't because of good technology, but because of soft porn.

Junjie Yan: We've done analysis. What truly makes users stay is definitely not so-called soft porn. For example, our product STARFIELD — its core is providing users with a platform to exercise creativity and imagination.

We've spent a lot of time and effort ensuring content is more positive, continuously improving platform safety capabilities.

LatePost: How much can technology improvements boost a product? You used MiniMax's self-developed MoE model on STARFIELD. How were the results?

Junjie Yan: Message volume increased 40% on launch day. Response became faster — previously 4 seconds, now 1 second. This wasn't just because of MoE, but also some other inference optimizations.

LatePost: Is faster technology improvement and larger user base a causal relationship?

Junjie Yan: This is very tricky. If you're the industry leader, OpenAI, then it's probably causal. If you're not the leader, then it's not causal.

Over the past year, many Chinese large model companies haven't had many users, yet their technology still improves — because you can progress just by learning from the leader. But long-term, if you believe your model can approach the best models, then user weight and value become increasingly higher.

This is like compute. Does having more compute let you make better models? Not necessarily — improving data quality might be higher ROI. But long-term, with more compute, you can definitely make better models. So it depends on the timeframe.

LatePost: AI-native super products versus mobile internet era super products — what do you think will be different?

Junjie Yan: When building mobile internet products, people cared deeply about whether they found a user pain point. But last year, the six or seven AI-native products with over a million DAU weren't designed around pain points — they released breakthrough technology and gradually became products. Conversely, when they later designed features more deliberately, they weren't as successful — like ChatGPT Plugins and GPT-S. If technology progress slows down, it will become product-driven again.

The current product approach is still technology-driven, not product-driven.

LatePost: Your product features are already quite granular now. For example, Hailuo AI frequently sends push notifications to attract users to open the app. You've actually done quite a lot of product optimization?

Junjie Yan: Recently we've also been reflecting — having too comprehensive product features might be a somewhat negative thing. It suggests you haven't spent the most effort on your most core functionality.

LatePost: What goals have you set for the team this year?

Junjie Yan: Technically, how to reach GPT-4. On the product side, how to 10x our user base, with a single product breaking through 10 million DAU.

LatePost: 10x growth — that's massive.

Junjie Yan: Actually, it's not that big. Mobile internet products all operate at the hundred-million DAU level.

You Can't Kill Competitors With Funding Alone

LatePost: How many AGI startups do you think China's current market capital and resources can support?

Junjie Yan: It won't be just one. The total resource pool is sufficient.

LatePost: Many investors have already stopped looking at large models. They believe startups have no chance in this space.

Junjie Yan: I lived through the last AI development phase that was built on stacking up funding rounds. If a company needs to keep raising money to survive, its real optimization becomes figuring out how to convince investors to give it more money.

My own internal path is to gradually serve users and generate some reasonable commercialization. Of course, with massive R&D investment, this is hard to achieve in the short term, but I believe we should explore this path.

LatePost: When overall market resources are limited, shouldn't the number one player try to raise the most money and starve everyone else? A lot of the last mobile internet competition worked that way.

Junjie Yan: The idea that you raise money frantically and others won't be able to — I don't think that's right. You can't kill competitors with funding alone.

Because among the leading Chinese startups, no one has an order of magnitude more resources than anyone else. The inflection point can only come from leading in technology, product, or commercialization efficiency.

LatePost: What about compute? That's scarce too.

Junjie Yan: China has compute now, more than before. Also, it comes back to Scaling Laws. When compute is insufficient, you need to find ways to optimize Scaling Laws by several multiples to achieve similar results.

LatePost: How do you assess the gap between you and OpenAI?

Junjie Yan: We have our own metric — you could call it "out-of-the-box usability rate." It looks at whether a customer or developer can quickly fulfill a complex need after plugging into a large model API.

From our own open platform's perspective, GPT-4 can handle almost any request. For example, one request we encountered last year: a user provides a novel and asks the model to generate a multi-character, voiced audio drama with distinct tones.

Very fine-grained use of GPT-4 could do this. Our own model couldn't at the time, but now it can.

LatePost: What about the gap with your Chinese peers?

Junjie Yan: We haven't tested against all of them. Because testing or not testing won't change what we do.

LatePost: What will happen in China's large model industry in 2024?

Junjie Yan: Chinese companies will produce something comparable to GPT-4, and more than one will do it. But the bigger question to think about is: what comes after that?

Treating the Company as a Function

LatePost: You said what's written in papers isn't a moat. Then what is the real moat in this space?

Junjie Yan: It's quite remarkable when you think about it — Pinduoduo started as Pinhaohuo, Meituan started as group buying, ByteDance started with Toutiao. None of them began with the products that eventually made them huge.

The difference between moderate success and massive success is that the massively successful companies all innovated organizationally, which enabled them to keep producing increasingly powerful things.

LatePost: Isn't the moat the people who write the papers?

Junjie Yan: I'll say something pretty scary: among the top 20, even top 50 people who've contributed to this field, probably not a single one works at a Chinese company right now.

The genius path doesn't work for us currently. The only viable approach is to gather people with sufficiently strong fundamentals, build a good growth-oriented organization, keep breaking through challenges together, and let everyone grow rapidly. I hope that in three years, the top 20 or 50 contributors to this field will come from Chinese companies.

LatePost: How do you want to build this organization?

Junjie Yan: I see it as optimizing a function. This function has no analytical solution — the essence is finding the direction of steepest gradient descent.

LatePost: Can you give an example? How do you find the direction of steepest gradient descent?

Junjie Yan: For example, in accelerating technological progress, it's learning from OpenAI, because that's the most certain path.

I don't mean matching their model parameters, but learning how to make experimental methodology more scientific; how to fail faster and iterate more efficiently; how to define problems more clearly and concisely.

LatePost: Pursuing gradient descent can trap you in local optima and miss long-term targets. How do you avoid that?

Junjie Yan: Our own evolution has been from looking at data very vaguely, to looking very deeply at data, to realizing that data alone isn't enough — you need to add better insight.

A lot of insight actually comes from long-term thinking. For example, if you only look at short-term product data, you wouldn't realize you need to build a new multimodal model.

LatePost: But can optimization function methods handle human problems? Like tension between tech and product teams?

Junjie Yan: When designing experiments or products, we make data instrumentation more granular, and use these data points to infer the real problems as much as possible, rather than relying on my or anyone else's subjective judgment.

We believe in data science. We didn't invent these things — Chinese internet companies have already done them extremely well.

LatePost: You previously said you wanted the organization to be lighter, but you're already at 300 people, most hired in the past year.

Junjie Yan: It's actually still very simple — only three layers in the organizational structure: me, my direct reports, and their direct reports.

You could say we have only three departments: one technology department, which I lead; one product department, split between C-end products and the open platform, each with one head; and one operations and growth department, which handles both product growth and company growth, with HR also under it, all with one overall head.

LatePost: Your peers — Zhipu AI has about 1,000 people, Moonshot AI has about 200, you're at 300. What lies behind these headcount differences?

Junjie Yan: It depends on what you believe. We don't need to prove anything to others — we just believe in what we're doing. Some unnecessary roles, we simply don't need. Whatever we need to do, we hire people to do that.

But we do need to build frontend products at a certain scale, so beyond algorithm and applied data talent, we also need people for inference systems, online services, development, and product operations.

LatePost: What talent do you need most right now?

Junjie Yan: More algorithm people. We now know how to run experiments, and our resources allow us to run many experiments, but we don't have enough people to run them.

Video generation models will become very practical this year. Based on last year, the first product to market had a much better chance of major success. Now many companies are racing to be first.

LatePost: How do you identify people who fit you?

Junjie Yan: Their joining raises the team's overall output. But this requires some posterior verification — some very strong people actually can't integrate into the team, while some who seem less strong can make the overall output stronger.

So in interviews, I pay attention to how they collaborated with people around them on important projects — with their mentor, with upstream and downstream colleagues.

LatePost: You managed a large tech team at SenseTime. What have you learned about managing technical people?

Junjie Yan: When you start wanting to do management, you may already be going off track.

What matters most is how to get everyone to build stronger things together, exceeding user expectations and the team's own expectations. AI may be a hot industry right now, but it's not that magical. At minimum it's a science, so do things scientifically: first, overall talent quality needs to be high; second, the organization needs a data-science-like methodology to quickly identify what works.

These two things combined — that's what we really need to do.

LatePost: How do you attract stronger people to join you?

Junjie Yan: Fundamentally, the organization needs to be strong and capable of continuously doing good things. This is the only path we can find.

LatePost: What culture do you want the company to develop?

Junjie Yan: First, no shortcuts — we've taken shortcuts many times and gotten badly burned each time. Second, User-in-the-Loop. Third, technology-driven.

These are all summarized from our previous experiences and lessons.

I Feel Like I'm Slowly Becoming a Set of Basis Functions

LatePost: SenseTime was your first job. What mark did it leave on you?

Junjie Yan: Mainly confidence in the technical path of concentrating forces to accomplish big things.

Some feedback was also searing — that's why I want MiniMax's organization to be simple enough. When people in an organization feel something is wrong but don't say it directly, that's hugely damaging to everyone.

LatePost: AGI was still non-consensus back then. How did you realize it was the direction?

Junjie Yan: It actually came from a moment of accidental reflection. In 2020, I was still leading a tech team at SenseTime. One day I suddenly realized I could no longer finish reading all the daily papers in AI — that struck me deeply.

As someone in technology, the daily technical progress had already exceeded my comprehension. Human evolution is very slow. The only way is to have better artificial intelligence to help technology develop, or to accelerate human research speed.

I had another observation at the time: before 2020, AI — much of what I did at SenseTime — didn't bring that much benefit and value to society.

That creates a huge contradiction: you believe AI will create long-term value for society, and that it's the only way to accelerate human technological progress; yet much of what you were doing wasn't directly contributing to that.

Was it because people didn't take it seriously enough? Clearly not — society was paying enormous attention to AI and pouring massive funding into it. Given all that, the only possibility was that our technical approach was wrong, or that the problems we were focused on weren't the ones AI should actually be solving.

LatePost: Many in the previous generation of AI practitioners actually recognized this contradiction, but none of them could find a way out.

Junjie Yan: OpenAI's release of CLIP in early 2021 was pivotal for me. That's when I began to realize there was no fundamental difference between natural language and computer vision — they were just one unified machine learning system. I saw the technical possibility of more general-purpose AI emerging.

When something like that happens, if you truly believe in AI, you have to go do something about it.

LatePost: How did you learn?

Junjie Yan: Surround yourself with people better than you. That's one of the few short-term satisfactions entrepreneurship has given me. I was fortunate to meet some truly top-tier people who gave me perspectives I wouldn't have had otherwise. When you think from a higher level, a lot of things actually become less difficult. Second, I read a lot of papers.

LatePost: You said to avoid being well-rounded in product. Are you well-rounded yourself? You rose quickly at SenseTime, from R&D all the way to Group VP — it seems like you could handle any function.

Junjie Yan: I don't think I'm well-rounded. The fact that I could do many different jobs probably has to do with my upbringing. I was born in a small county in Henan; there was no one around to teach me anything, so I had to figure everything out myself. That developed my ability to intuit how things work. I didn't want to be this way — I was forced into it.

But looking back today, that ability has proven extremely useful. When I take on something I've never done before, I can quickly find the underlying logic.

LatePost: What do you see as your weaknesses?

Junjie Yan: Although I've done technical work, I'm not a top-tier researcher. Maybe just second-rate.

LatePost: That's hardly fair — your papers have nearly 30,000 citations on Google Scholar.

Junjie Yan: The top person in the world probably has 300,000.

LatePost: You said to treat the company as a function. What kind of function are you?

Junjie Yan: (long pause) Back in school I learned about Taylor expansion — how a complex thing can be approximated by combining simple functions.

In other words, you can use a set of basis functions to approximate any function. I feel like I've gradually become a set of basis functions, combining in different weights to take different forms as needed.

LatePost: We've talked for so long and haven't touched on changing the world, changing humanity.

Junjie Yan: The things you truly want to do shouldn't be talked about every day.

LatePost: Can we talk about it today?

Junjie Yan: Still "Intelligence with everyone." There are two meanings to this: one, we want to serve every person with the best technology; two, in our journey toward AGI, we need to iterate and grow together with our users.

And I've seen technological progress happening faster than I imagined.

Qianming He also contributed to this article.

Cover image source: Ford v Ferrari