Li Shanyou in Conversation with Yachen Song: Building the World's Most Powerful AI 3D Foundation Model in Three Years | Oasis Vitality
"Everyone has their own happiness, and that matters a lot to me."

Oasis Capital was the sole lead investor in VAST's angel round.
Since our investment in January 2023, we have continued to follow and believe in VAST's long-term commitment and execution in the AI 3D direction.
Two years ago, at the most crowded stage of large model competition, founder Yachen Song chose a path with lower consensus but greater systemic challenges, and has consistently pushed it toward verifiable product and commercial form.
At the end of this year, Yachen Song was invited to participate in Shanyou Exploration Flow, a video podcast produced by Hundun Academy and hosted by Professor Li Shanyou, where he engaged in a deeper exchange around technology choices, creative tools, and his unique way of thinking about the world.
Below is the full transcript. Enjoy.
He is the first — and to date, only — Chinese person to take the keynote stage in SIGGRAPH's 50-year history, sharing the platform with industry leaders like Jensen Huang of NVIDIA.
He was MiniMax's Employee No. 001. Just as large language models were surging, he turned around and plunged into the uncharted territory of AI 3D, an adventurer set on regenerating the three-dimensional world from scratch.
He is the entrepreneur who, in just two years, led his team through three rounds (each worth hundreds of millions of RMB) of funding, with a valuation firmly at the top of global AI 3D large model companies.
What he aims to do sounds crazy but beautiful — a 3D version of TikTok.
He is Yachen Song (Simon), founder and CEO of VAST, a student of Hundun Academy's 6th cohort. He wants to build the world's leading AI 3D large model.
This young entrepreneur, born in 1997, has in less than two years led his team to push the Tripo series of 3D large models from technical prototype into the hands of tens of millions of users: 8-second text/image-to-3D generation, the first to validate the 3D Scaling Law, parameters soaring to 20 billion — thrusting generative 3D AI directly into the "IMAX era."
While technology surged forward, commercialization kept equal pace. As of August 2025, VAST's annual recurring revenue (ARR) reached $12 million — industry-leading. Around 5 million professional users, over 80% from overseas. On the B2B side, more than 40,000 SMEs and approximately 700 large enterprises use their tools.
But none of these "hard metrics" are the most fascinating thing about this person.
The real contrast lies in this — with the most hardcore technology, he thinks about the most ancient proposition: how to maximize the sum of human happiness. He earned a double bachelor's degree in International Relations and Economics from Johns Hopkins University, while also passionately studying theology.
This is why, in the large model track where it was easiest to "ride the momentum," he deliberately turned to do something harder that almost no one dared: letting AI not just generate video, but regenerate the "three-dimensional world."
Professor Shanyou said: "My good friend Jinjian Zhang of Oasis Capital mentioned him to me with one word: 'little monster.' He said you absolutely must meet this kid — there's a very rare vitality in him. So today I've invited him to my podcast."
He is an "idea-driven entrepreneur," the kind you look at once and know he will rebuild the world. In this conversation, you will hear:
- His penetrating understanding of AI 3D, a rare and brilliant analysis within the industry
- How he views the essential relationship between technology and "human experience"
- His grand vision: everyone living in real-time interaction within "the world they love most"
If you want to meet a young founder who truly "enters the Way through commerce," if you want to truly see how a young entrepreneur evolves, leaps, and awakens across disciplines — this conversation is a must-listen in full. The knowledge density is extraordinarily high.
Welcome to the full episode. Step into the world of this pure and powerful "little monster."


Origin: "Let Creative People Focus on Creativity Itself"
Li Shanyou: I first learned about you two years ago.
Jinjian Zhang said we should accompany some "little monsters" as they grow. He said some young people have their own independent ideas, always positive and upward, yet feel lonely in their environment. His words deeply moved me, and that was when I first knew of you. Please briefly introduce your entrepreneurial journey.
Yachen Song: I joined SenseTime in 2018, where I assisted Li Xu with various tasks (Li Xu is co-founder and CEO of SenseTime, also a computer vision scientist). I chose the AI plus animation field partly because I personally have strong interest in games and animation and find the domain fascinating; partly because I discovered a core problem through strategic analysis. People typically assume animation, film, games and such are creative industries, and the logic seems simple: whoever is smartest and most creative stands out. However, I found that whether in China or globally, the animation industry is actually labor-intensive, more like "turning screws" work.
If you recruit animators, most resumes come from graduates of the eight major fine arts academies. These art school graduates should be brimming with creativity, yet in their work they often end up doing repetitive tasks — converting models to 3D, or adjusting animation frame by frame to make a character run. This work is like turning screws, highly mechanical and repetitive.
This model prevents the animation industry from truly becoming a creative industry. Because those with creativity, after long-term repetitive work, gradually have their creativity worn away. By the time they finally become producers, little creativity remains. This is one reason why China struggles to produce quality animation content.
I believe the animation industry shouldn't be labor-intensive, but a true creative industry. Based on this view, an important question we faced was: how can this industry return to its creative essence and achieve industrial upgrading? The answer is to accelerate the introduction of AI technology. Because AI can release people from repetitive work, allowing creative talent to focus on creativity itself, thereby driving innovation and development across the entire industry.
Li Shanyou: Let AI handle those repetitive, labor-intensive tasks, while letting creative people focus on creativity itself.
Yachen Song: Yes. That was the significance of AI at the time. So we began discussing the concept of AIGC (AI Generated Content) very early on.
Li Shanyou: What year did you start this?
Yachen Song: It should have been the second half of 2019. We discovered a problem: AI technology at the time was not yet able to solve these issues well.
Specifically, while providing services to many Chinese animation companies, we encountered challenges on two fronts. On one hand, the commercialization level of China's animation industry was low, and these companies themselves had limited funds, so we couldn't generate sufficient revenue from them either. On the other hand, the technology at the time was not mature enough to truly help them unleash creativity and solve the problem of repetitive labor.
Nevertheless, I realized that even if the technology was not yet perfect, we could still begin applying AI technology. So I personally took on the roles of director and screenwriter, responsible for content creation and IP design. Starting from zero, we gradually built various IPs from hundreds of thousands to millions of followers. This process was deeply meaningful to me because I inherently love content creation and creative work. These new IP creative contents were mainly presented in short video form.
At the time, Li Xu was very supportive of my ideas. We also assembled an animation team of about forty to fifty people, attempting to fully apply AI technology within the company to produce animation. Unfortunately, the profit margin was limited, forcing us to reconsider our direction.
So we began searching for more profitable domains. At that time, the gaming industry was exploding — games like Romance of the Three Kingdoms: Strategic Edition and Genshin Impact were hugely popular. Especially when games combined with concepts like the metaverse and AI, the industry developed rapidly.
Based on this market understanding, we integrated our existing AI technology into gaming solutions and began promoting AI technology in the gaming industry. Later, with the rise of the metaverse concept, business scale continued to expand.
However, I left SenseTime around June-July 2021, and then formally participated in founding MiniMax. At the end of 2022, I left MiniMax. One important reason for leaving was that I believed the industry's blind rush into AGI or language large models at the time was an emotional product — everyone was trying to become the next OpenAI, but this herd mentality was not rational.
Additionally, I observed that from 3D to video to images to text, information density gradually decreases — this is a compression process where information is progressively lost. We ourselves live in a 3D world. When a child is born, if you give him a ball, he will instinctively interact with it — this interaction is natural. However, text, images, and video are more common in the internet era because internet technology was not yet mature enough. In fact, the world was not originally dominated by text, images, and video. We are more interested in artifacts with writing because text has low information density — a small amount of text can abstract many things.
But in prehistoric civilizations, most things were geometric sculptures — tools, ornaments, totems — these were the mainstream forms of expression at the time. As humanity developed, people discovered pigments and began painting in caves. This form of expression has even lower information density, but can more vividly express more content. Later, text gradually emerged.
By the same logic, in the internet era, with limited bandwidth and processing power, information dissemination began with text (such as Weibo, blogs), gradually developed to text-and-images (such as WeChat Official Accounts, Xiaohongshu), then to video (such as Douyin, Kuaishou, TikTok). As internet technology matures, information dissemination should be a gradual "decompression" process, returning to the most authentic state. People no longer need to consume compressed information, but directly consume the most authentic content.
When training AI or developing general large models, using native data with the highest information density is clearly more valuable than using compressed information. Because native data contains greater information volume and is closer to the essence of things.

Insight: "3D Is the Nature of the World" and a Technological Gamble
Li Shanyou: So you don't believe that "language is the inevitable path to AGI"?
Yachen Song: I believe that understanding the world and expressing it in 3D carries the greatest information density. 3D is the most sincere, most authentic, and most reliable information carrier and content medium. We live in a 3D world. If you try to compress the world's information — say, through video — you get two approaches: live-action filming and virtual production. Live-action means choosing a position and angle in the physical world to shoot from. Virtual production means shooting in a human-created virtual world, like animated films such as Ne Zha and Avatar.
Both approaches share one thing in common: they each have a real or virtual 3D world as their foundation.
Now a new approach has emerged, called video generation. The problem with this method is that it tries to deceive the audience, because it doesn't have a real 3D world as its foundation. This approach is inherently distorted. When it attempts to construct a false world, it creates countless lies that need to be covered up.
For example, it produces consistency issues and memory duration problems. Suppose there's a cup in the video. With normal live-action filming, if the crew goes out to shoot for five hours and comes back, the cup will definitely still be there. But with video generation, if it generates five hours of video, it might forget the cup ever existed. These hallucinations, consistency failures, and memory problems all stem from the fact that video generation is lying — it's not real. The same problems appear in image generation and text generation. They're simply making things up, which creates fundamental issues.
So we say 3D is the universal solution. Through 3D, the most generalizable form, we can provide the most information to train AI. Once AI is ready, whether it's AI-generated content or the judgments it makes, everything can then be compressed. Content produced this way would be the most sincere and authentic.
For instance, to solve the memory duration problem in video generation, you could place a marker in the 3D world — quietly telling the AI that there's a cup here — so when it returns, it remembers the cup exists. This marker can be expressed in different ways, through cinematic expression or display expression. For example, through a beam of light or a QR code, the machine can calculate and discover there's a cup here, then display it. This requires a genuine 3D expression to solve the falseness in generated content. It's just a different form of expression.
In short, 3D is the most authentic and information-rich; it most closely matches how the world originally is. Training, adjusting, and developing on a 3D foundation — whether for AGI or anything else — this is the nature of the world. Otherwise, building new content on a foundation of lies only produces more lies, ultimately leading to all sorts of hallucinations and problems.
Li Shanyou: Very well said, truly excellent! Looking at actual development, the mainstream trend right now is indeed starting with text. Because text was the earliest to be used for training language models, and then large language models gradually emerged from that.
So many people believe language is the inevitable path to AGI. But what you said makes a lot of sense — language is essentially a compressed information carrier, while the 3D world is what's closest to reality, the least compressed information source. The 3D world contains rich, uncompressed information; this is what produces genuine knowledge and models. Your thinking is excellent.
What I want to ask is: when you first started your company, or even before that, did you first have this idea about 3D and then decide to pursue it? Or did you see others working on 3D-related things and decide to follow? In other words, was your decision based on your own independent understanding, or on observing and learning from others' experience?
Yachen Song: First, we're definitely 3D-based. We believe 3D is extremely valuable, especially the interactivity that 3D enables — this real-time interactivity is something other forms lack. We had an important realization at the time: from text to images to video, these content formats are essentially not real-time interactive.
While technically not completely non-interactive, people typically don't use these formats for real-time interaction. So we call text, image, and video content a type of experience, one that could be called "empathy" — experiencing by "standing in someone else's shoes." For example, when you watch the TV series The Knockout, you experience Gao Qiqiang's life; when you read a novel, you experience Zhang Wuji's life. This content lets you gain experience by observing others' stories, rather than directly participating yourself.
Li Shanyou: 3D means we're not just bystanders.
Yachen Song: Exactly. In the 3D world, the core is "subjectivity" — the "I" here is an autonomous, agentic being. For example, I can have the overpowered ability of "999 damage per strike," go out and conquer the world — this is entirely an immersive, first-hand experience centered around "me." This experience is fundamentally different from other types, and right now, these subjectivity-centered experiences are actually still quite scarce.
Li Shanyou: But when I play 2D games, isn't that me experiencing it too?
Yachen Song: The 3D format itself is the most suitable vehicle for interactive experiences — after all, humans are naturally accustomed to interacting with their surroundings and other people in three-dimensional space. This innate behavioral logic creates an extremely strong binding relationship between 3D and "interactivity." Because of this, in current understanding, when we see the concept "3D," it almost by default implies interactive properties.
Li Shanyou: 3D equals interactive.
Yachen Song: The industry is already moving in this direction, and this has become an established fact. When we experience various content from an empathetic perspective, we can clearly see this type of experience is already quite abundant. In our daily lives, we have access to social and short-video platforms like Weibo, Xiaohongshu, Douyin, and TikTok, as well as long-video platforms like Netflix and iQIYI — watchable, relatable content is everywhere, satisfying our emotional resonance needs in many ways. But in contrast, experiences centered on "subjectivity" are remarkably scarce: in the physical world, we can autonomously make choices and direct our own actions, so this type of first-hand experience is relatively rich; but in the virtual world, this kind of "me"-controlled, self-directed experience is still in a very impoverished state.
Li Shanyou: That is indeed the case. Why does this phenomenon exist?
Yachen Song: The reason is that text, images, and video have already attracted mass participation in creation — this is what's called UGC (user-generated content). But 3D or interactive content still belongs to the "elite" art form; this is the most fundamental difference. In the past, text content was extremely scarce. For example, in the Tang Dynasty, the number of people who could write poetry or novels probably didn't exceed one million — this was elite art. The same was true for images; in the past, what we saw in galleries were mostly works by masters like Michelangelo, and worldwide there were fewer than one million people who could create such works. Video was similar too — whether Hollywood or Hengdian, the number of people engaged in professional video creation was also fewer than one million.
Today's 3D or interactive content is the same. Companies like Tencent, NetEase, and Ubisoft — globally, the number of professionals who can engage in this type of creation probably doesn't exceed one million either. So how do we get the masses to participate in 3D or interactive content creation?
The key is having a mass-market creator tool. For example, text creation has typing methods; image and video creation have smartphone cameras — these tools let everyone create content with zero barrier, zero cost, in real time. Why must it be zero barrier, zero cost, real-time creation? Because the biggest difference between UGC and PGC (professionally generated content), and between the masses and professional users, is: professional users do it for money, while UGC users don't. This is the most fundamental difference.
Li Shanyou: It's about expression, about entertainment.

Execution: 3D TikTok, Finding Everyone's Optimal Experience in the Moment
Yachen Song: Users' motivation for participating in creation isn't profit to begin with — it's more about emotional expression, like showing off, venting, sharing snippets of their lives. So first and foremost, we need to ensure they "don't lose money" when creating, that there's no financial burden.
So the question becomes: how do we truly get the masses to participate? This must satisfy the core need of "zero barrier, zero cost, real-time creation."
We've noticed that AI 3D large models happen to provide this possibility: they have the potential to turn interactive content or 3D content creation into a mass-market tool that anyone can use, letting every ordinary person easily participate in creation. When creation barriers are completely broken down, large numbers of users will flood in and produce content, creating a reverse cycle: first, the proliferation of creation tools brings an explosion of content, and this massive volume of content needs a dedicated platform to carry and distribute it, ultimately giving rise to something like a "3D TikTok" product, or a 3D UGC-centered ecosystem.
Once such a 3D UGC ecosystem takes shape, the quantity and variety of interactive content will experience explosive growth, and the entire interactive world will become incredibly rich. Just imagine: when everyone can freely choose from infinite virtual worlds or interactive content to personally experience, isn't that, in a sense, bringing heaven to reality? Because everyone can, in the present moment, find and experience what is truly optimal and most extreme for themselves.
Li Shanyou: You're absolutely right — understanding ultimately needs to land in practice. Behind this, two tracks are actually advancing in parallel: on one hand, understanding needs information to support it and practice to ground it, and your reasoning at the understanding level just now was truly excellent. Now let's pull back to the practical level: how did this understanding translate into concrete action? Was it because you saw certain clear signals that led you to do this? Or when you started, was this field completely blank in the world? How did you initially get this started?
Yachen Song: I'm definitely not the only person who thought of this — many people in the world have seen this direction and are already working toward it. We previously quietly built a product similar to a 3D TikTok, but then discovered a problem: once the product reached a certain stage, content creation became very difficult to keep growing. We did extensive user research, and after talking with people, we found the core crux was that users need a zero-barrier, zero-cost creation experience — only then are they willing to actively participate.
So we realized we had to build a mass-market creation tool first. After that, we started looking for the right technical path, and discovered that AI 3D was already showing the first glimmers of dawn — it had become viable for real-world deployment. So we focused our energy on seriously refining the AI 3D technology and product, and that's precisely how we arrived at where we are today.
Li Shanyou: What's the core difference between this kind of 3D foundation model and the language models we're familiar with? When you first started out, did you begin by building the 3D foundation model, or did you develop the front-end creation tool first?
Yachen Song: We started with the foundation model. The tool is something we only began working on this year.
Li Shanyou: You focused on foundation models from the very beginning of your founding, specifically pushing forward 3D foundation models. This typically requires enormous conviction and foresight, because most companies would choose to develop tools first.
Yachen Song: Yes.
Li Shanyou: Let's talk about 3D foundation models.
Yachen Song: I believe "foundation models" actually represent a paradigm shift in thinking. Why do I say this? In the AI 1.0 era, the mainstream approach in the industry wasn't about pursuing model scale — on the contrary, it was about making models as "small" as possible. The R&D logic at the time was relatively straightforward: assemble top algorithm scientists, tackle specific and long-tail scenario problems one by one — like facial recognition, abnormal behavior detection — and through extensive manual parameter tuning and training, build the lightest possible specialized models. The smaller the model, the lower the training and deployment costs, and the clearer the commercial returns. So the core of that stage was competing over who could make models smaller and more efficient when solving specific problems.
By the AI 2.0 era, thinking had fundamentally changed. People began exploring: could we use massive data and powerful compute to drive the construction of an extremely large and general-purpose model, one that could generalize to almost all scenarios and solve in one go the problems that previously required countless small models to cover? This follows the famous Scaling Law. Just as GDP in economics depends on labor and capital, in AI, model performance can be viewed as a function of data and compute. When both grow in tandem, performance improves significantly; but if only one grows while the other stagnates, marginal returns rapidly diminish. It's like having millions of workers but only one shovel, or millions of shovels but only one worker — either way, efficiency can't improve.
We're currently in the middle of this paradigm: compute is still growing rapidly, but the supply of high-quality data is gradually hitting bottlenecks, causing the marginal returns of compute growth to decline. So the industry has also begun reflecting: does this mean we need to return, to some extent, to the AI 1.0 approach, once again leveraging lighter, more focused models to solve specific problems? There's no definitive answer yet, but what's clear is that these two modes of thinking are forming a beneficial complementarity and cycle.
As for the difference between language foundation models and 3D foundation models, I think it manifests more in technical path and domain transfer. When a breakthrough technology emerges — say, the Transformer — its core ideas often cross domains, inspiring scholars in other fields to think: "Could my field adopt this paradigm too?" This kind of cross-domain technical borrowing and conceptual migration is precisely what drives progress forward.
Whether it's Diffusion, Transformer, or "foundation models" themselves, their core value lies not merely in the specific technologies, but rather in the general problem-solving paradigm they represent.
Li Shanyou: But from an outside perspective, language foundation models are already complex enough, and 3D foundation models are generally considered to present even greater technical challenges.
Yachen Song: The difficulties mainly come from several things. First is the scarcity of interdisciplinary talent. Building 3D foundation models requires deep integration of expertise across three domains: artificial intelligence, computer vision, and graphics. This means teams must be proficient in distributed training and parallel computing required for large models, deeply understand low-level visual information processing, and also master complex geometric representation and rendering techniques in graphics. Such cross-disciplinary top talent has always been extremely rare in the market — this is essentially a brand new field with virtually no ready-made senior experts. Therefore, team building often has to start from foundational cultivation, or rely on young talent with learning ability and cross-disciplinary backgrounds.
Second is the severe shortage of high-quality 3D data. As mentioned earlier, due to limitations in internet ecosystems and terminal devices, what humans have long consumed are essentially "compressed packages" of 3D information — namely text, images, video, and even livestreams. These are all two-dimensional carriers that represent drastic simplifications and projections of the three-dimensional world. The native, structured, large-scale 3D data that we truly consume directly and can use for model training — such as detailed models, point clouds, dynamic scenes — is exceedingly scarce. This fundamental lack of data constrains the development and training effectiveness of 3D foundation models.
The third problem is that the 3D field lacked many resources in its early days, so its development speed was inevitably limited.
This shift has been particularly evident in computer vision. Take CVPR, the top conference in computer vision: early on, a large number of computer vision-related papers emerged at SIGGRAPH, the top graphics conference, even "encroaching" on some of SIGGRAPH's content. This was because talent researching graphics and AI 3D was relatively scarce, almost negligible. Yet in just two short years, the situation changed dramatically. Now, not only has SIGGRAPH itself seen a surge of content related to AI, 3D, and graphics, but CVPR has also produced numerous outstanding papers related to AI, 3D, or graphics, including best paper awards and other major recognitions.
This shift shows that as AI, 3D, and related fields have gradually become prominent disciplines, attracting substantial resources and capital investment, a virtuous cycle has formed. Looking back at the early days, the field faced many challenges: first, a lack of professional talent; second, insufficient data resources; and finally, because it wasn't yet a prominent discipline at the time, overall resources were scarce.
Li Shanyou: At that time, you had neither relevant technical background nor much funding as a startup, and this was a difficult undertaking. How did you pull it off?
Yachen Song: Mainly by asking others for advice and finding partners.
Our CTO, Ding Liang, gave me a lot of guidance. We were colleagues at SenseTime, and we trust each other. I have great confidence in him and the team's technical capabilities — I believed they could deliver on the technology side. Later, Chief Scientist Yanpei Cao and other young scientists joined one after another. Our technical team is very strong, and I trust them deeply, so I focused my energy more on data, resources, and other matters. We built a capable technical team in a short time, and I could confidently hand over the technical work to them.
Li Shanyou: When you first started your company, were there any 3D foundation models globally?
Yachen Song: There were probably some related papers, especially from overseas — early attempts by Facebook, Google, NVIDIA, OpenAI, and others — but they couldn't really be called AI 3D foundation models. There weren't truly foundation models in the real sense.
Li Shanyou: This is quite interesting. At SenseTime, you worked on AI-related tasks, mainly finding various application scenarios, and at MiniMax you encountered foundation models. If you were to start a company, the easiest path would be to enter various vertical domains based on foundation models, like SenseTime did. But you didn't choose this path — instead you went one level deeper. This is a philosophical kind of entrepreneurship. Where did your confidence come from? Your starting point was very unusual. Why did you have such confidence? Was it just some inexplicable force?
Yachen Song: I think if I were doing this alone, I definitely would have felt it wouldn't work out. But I firmly believed we had an excellent team, and my trust in the team was very strong — I never doubted that the team couldn't achieve our goals.

Business: "This exceeded expectations, faster than I imagined"
Li Shanyou: Was the startup idea yours or the CTO's?
Yachen Song: I proposed the startup first.
Li Shanyou: Then in the beginning, what drove you? What motivated you to do this?
Yachen Song: We genuinely felt there was this need. It's like we're trying to reach a certain goal, and along the way we encountered a nail — we need to find a hammer. What kind of hammer is suitable? We felt this hammer was the most suitable one. This is indeed different from other companies: many AI companies first build a hammer, and then because certain things become hot, everyone says, now that we have this hammer, let's go find application scenarios, find nails. But we actually encountered a nail during our entrepreneurial process — creators have no way to create in real-time with zero barrier and zero cost — and we needed to find a hammer to solve this problem, and this thing is the best hammer.
Li Shanyou: Can I understand it this way: the 3D TikTok idea came first.
Yachen Song: It's actually a vision, but to realize this vision, you might need to hammer a nail first.
Li Shanyou: Right, people come to create, they need tools, and the tools require a 3D foundation model. So you derived this step by step. But ultimately what you want to do is a 3D content creation platform, similar to 3D TikTok. From the demand side, the scenario side, you're clear, and based on this demand you derived your way to this point.
Yachen Song: I think 3D TikTok, or an interactive content platform, is definitely a long-term need. Even if I don't build it today, someone else will certainly build it tomorrow — this is a consensus.
Li Shanyou: So over these past few years, how has your 3D foundation model developed?
Yachen Song: I think the development speed has been faster than I imagined.
Li Shanyou: Why?
Yachen Song: Probably because my previous experience was in the AI 1.0 era, when technology didn't develop this fast. You'll find that the pace of technological development in the last two or three years has been somewhat "abnormal" — people have become numb to it. Actually, the technological development of the past two to three years has been highly abnormal, the speed is too fast. People have seen too many spectacles, to the point where they've become numb to real technological progress.
Li Shanyou: Now it's exponential progress, and people feel indifferent about it.
Yachen Song: Take video generation, for example. If you put it 100 years ago, it would absolutely be a great invention, possibly the greatest invention of the century. But placed today, it's just one of many inventions that feels pretty decent.
This is something I find truly remarkable — it actually exceeds my comprehension. I originally thought that maybe in four or five years, AI 3D foundation models could enter the pipeline (the 3D pipeline is how we use computer language to express a three-dimensional world), and even surpass human levels, and that would already be quite good. But now, in just two or three short years, it's basically already achieved that. I think this exceeded my expectations — faster than I imagined.
Li Shanyou: Overall, on the user scenario side, what stage are you at?
Yachen Song: We currently have around five million professional users on our professional tools, with over 80% coming from overseas. We also do some B2B work, serving roughly 40,000+ SMBs and about 700+ large enterprises.
When it comes to use cases, we mainly have four categories. The first is content creation — things like games, animation, film and television, short dramas, CG, and so on. The second is industrial design, covering light industry, heavy industry, flexible manufacturing, 3D printing, and the like. The third is exhibitions and displays, such as e-commerce, advertising, education, cultural tourism, museums, and related fields. The fourth is emerging industries — embodied intelligence simulation, digital twins, digital humans, AI + gaming, world models, spatial intelligence metaverse, XR + AI glasses, and so forth.
Li Shanyou: Is your biggest challenge right now on the technology side or the market side?
Yachen Song: I don't think the greatest difficulty is purely technical or market-driven. It's whether people, living in this era full of noise and temptation, have enough patience and resolve to see something through. Long-termism is indispensable if you want to build something of value, something relatively great. Take OpenAI — it took six years of quiet accumulation to reach where it is today.
Doing something worthwhile necessarily requires long-term buildup and persistence. Along the way, you face countless temptations and fears, and these constantly test your resolve and patience. Over the past two or three years, technology has advanced at breakneck speed, transformation across every field has accelerated, and people's tendency to pivot has intensified dramatically. Yet in such a rapidly changing era, maintaining a certain kind of "slowness" carries its own unique value.

Philosophy: Everyone has their own happiness, and that matters to me
Li Shanyou: From a long-termism perspective, what's the ultimate vision for this?
Yachen Song: The vision is to contribute to civilization for the world, and to create happiness for humanity.
Li Shanyou: I think you're the first student I've encountered who has an obsession with ideas, and who can embrace the complexity of the world. Have you developed your own distinctive way of thinking?
Yachen Song: I think I do have my own distinctive way of thinking, but I'm not yet able to articulate it well.
Yachen Song: I think my thinking tends toward the theories proposed by Mill (John Stuart Mill) and Bentham (Jeremy Bentham). (These two are the main representatives of Utilitarianism, an important theory in traditional Western ethics that advocates pursuing "the greatest happiness.")
Here's how I understand and apply it: everyone has their own happiness.
A lot of philosophy is really discussing morality, while theology deals with questions like who is the First Principles, who created the world, where humans came from. When we talk about philosophy, we're actually discussing morality, but our way of thinking isn't based merely on these questions about origins and creation — it's more like a kind of thinking grounded in worldview and values.
In terms of thinking style, I believe the essence of morality should be maximizing the total sum of happiness. Take the trolley problem — it explains many issues in philosophical moral judgment quite well. Suppose there's a railway track: on one side, one person dies; on the other, two people die. I'd choose the side where one person dies, because that minimizes the reduction in total happiness. If one death is -1, then two deaths is -2. The calculation is quite simple.
Li Shanyou: So "maximizing the total sum of happiness" is an important phrase for you.
Yachen Song: Exactly. This also connects to what I do as an entrepreneur. For instance, I think there are three main directions for entrepreneurship — of course there are far more types than these, but these three are the most prominent right now. The first is characterized by rapid diffusion, like Elon Musk and Thomas Edison, who dedicated themselves to giving people more resources, like cars. The second is about helping people live longer — various medical companies aiming to extend human lifespan from 50 years to 100, to 1,000, even to immortality. But I prefer the third. Like The Walt Disney Company: even if there are only five people, and they only have three days to live, I still want those five people to live those three days as happily as possible. To me, that's what matters most.
Li Shanyou: Hmm. When you're doing this, what's most important to you? What's your first principles? Where's your core conviction? For example, Elon Musk says he wants to make humanity a multi-planetary species. That matters enormously to him — he feels his life would be wasted if he died before accomplishing it. But Jensen Huang definitely doesn't think that way. His first concern is survival.
Yachen Song: I think people can choose their own most extreme experience, and that matters to me. I even feel that everyone being able to have their own most extreme experience is something rare and precious.
Li Shanyou: That's what's most important to you, a belief you hold deeply.
Yachen Song: Yes, I think this is the most important thing.
Li Shanyou: Where lies your ability? Where's your talent? Why are you the one who can do this?
Yachen Song: I don't think it's about whether I can do it, but whether the direction is right. I can run slowly, so I'll run slowly. I can also accept that I might not be the one to ultimately accomplish this — it could be completed through collaboration with others, or someone else might do it in the end.
I'm unwilling to do something that seems like I'm good at it, but that I don't believe in or find meaningless. Conversely, I might not be skilled at making this particular thing happen. For example, I know nothing about technology, but I believe doing this thing itself is important. Whether I'm the most suited for it matters less.
Jack Ma probably wasn't necessarily the most skilled person to build Alibaba either. There might have been tens of thousands of people more capable than him at the time. But whether you do the thing — that may be the most important capability.
Li Shanyou: What you were trying to express just now is the significance of the thing itself. I think you're very fortunate, because you genuinely believe this matters to you. Not everyone can do that. You're an idea-driven entrepreneur, and you believe ideas matter to you.
Yachen Song: They matter enormously.
Li Shanyou: You're definitely in the minority. So I think you're a little monster — the kind I particularly admire, like, and am willing to accompany. Second, we've found this vehicle, and logically, it can lead to that goal.
I've been looking at Jensen Huang's life recently, and what moved me most is how his first half and second half differ. In the first half, he made gaming chips, full of competition, just struggling to survive. In the second half, he moved into GPUs, CUDA, accelerated computing, and artificial intelligence — I feel he was being himself. At that point, there should have been no competition. In the first half, his competitive strategy was non-competition; in the second half, he became himself.
I believe life has a first half and a second half. The first half is driven by EGO (the self), driven by greed, anger, and delusion, driven by human instinct. But I believe there is a second half, driven by truth, goodness, and beauty. This isn't just for me, or for everyone — it's for the universe. Most people complete the first half before entering the second half. But you're a rare exception. You've stepped directly into the second half. That's truly remarkable.
I think the happiest thing in life is doing what I was born to do, and within that, becoming the best version of myself. This is using the endeavor to cultivate the person — the unity of person and endeavor.
This is also what I mean by entering the Way through commerce.





