"Language Models Have Hit a Wall; 3D Foundation Models Are Just Getting Started" | A Conversation with VAST Founder Yachen Song on Two Years of "Wild Ride" in 3D Foundation Model Entrepreneurship
In the future, everyone will be able to create worlds with AI — not merely inhabit them.
In the future, everyone will be able to use AI to create worlds, not just use them.

👦🏻 Podcast interview: Koji, Ronghui
🥷 Editing: Starry
🧑🎨 Layout: NCon

Beyond the wave of text-to-image and text-to-video, where is the next AI technological singularity that will ignite our imagination? The answer may be AI + 3D.
This week, we invited Yachen Song (Simon), founder and CEO of 3D foundation model company VAST, to talk with us about the story behind VAST's latest 3D generative foundation model, Tripo 3.0.
This 1997-born entrepreneur has raised three consecutive funding rounds in short order — each in the tens of millions of dollars — building up enough ammunition. After a year of heads-down work, Simon is making his podcast debut this year to discuss several key strategic questions with us:
- He believes large language models have already "hit a wall," with evolution slowing down, which is what has created room for applications and agents to flourish. 3D foundation models, by contrast, are completely different — they've only just begun and remain a blue ocean.
- At the resource-constrained startup stage, why does VAST "want both"? Both developing foundation models and building its own application, Tripo Studio?
- Why is the ultimate form of technology a process of "decompression"? He argues that the history of human media (text → image → video → 3D) isn't about ascending dimensions, but rather a series of forced dimensional reductions and compressions of the 3D "source file" world due to technological limitations. Technological progress is the process of "decompressing" back to the world's true form.
- And in a future where robots can do everything for us, how will human value be redefined?
From attracting classmates to "top up" his hand-drawn paper RPG world with spicy strips in elementary school, to going all-in on AI entrepreneurship to build an "infinite world" where he believes everyone will create in 3D — welcome to Simon's observations and reflections on his entrepreneurial journey. We also welcome your thoughts on AI + 3D in the comments.

Listen on WeChat:
Listen on Xiaoyuzhou:


Rapid-Fire Q&A
👩🏻 Ronghui
Hello everyone, welcome to this episode of "Crossing." Our guest today is a young entrepreneur — Yachen Song, Simon, founder of VAST, a 3D foundation model research company. At the end of August, VAST released its latest 3D foundation model Tripo 3.0. Today we've invited him to talk about the story behind 3D foundation model development and some of his reflections from two years of entrepreneurship. Let's jump into rapid-fire Q&A.
Simon, please answer: Age?
👦🏻 Yachen Song
👩🏻 Ronghui
Alma mater?
👦🏻 Yachen Song
Johns Hopkins University for undergrad.
👩🏻 Ronghui
Your MBTI and zodiac sign?
👦🏻 Yachen Song
I'm ESTP, Pisces.
👩🏻 Ronghui
Describe your current company and product in one sentence.
👦🏻 Yachen Song
Our company is called VAST. We're an AI 3D foundation model company. We have a product called Tripo — the input can be text, images, or multimodal, and it outputs complete 3D content.
👩🏻 Ronghui
Funding status?
👦🏻 Yachen Song
We've raised three rounds previously, each around several tens of millions of dollars.
👩🏻 Ronghui
Team size?
👦🏻 Yachen Song
We're about 110+ people.
👩🏻 Ronghui
What were you doing before entrepreneurship?
👦🏻 Yachen Song
Entrepreneurship, actually. Earlier I worked at SenseTime for a while on AI + animation, AI + games. In 2021 I co-founded MiniMax, and in 2023 I founded VAST.
👩🏻 Ronghui
We have an icebreaker social segment — introduce 10 things about yourself using "I am XX." Simon, please give it a try.
👦🏻 Yachen Song
First, I am Yachen Song, 28, an entrepreneur.
Second, I am the founder and CEO of VAST.
Third, I am an addictive gamer, very obsessed with games. In college, I sat in my bed gaming so much that I left a dent in the mattress.
Fourth, I am someone who loves traveling. I've been to Georgia, Bermuda, Cuba, Turkey, Morocco, and other places.
Fifth, I consider myself a cross-industry person. My undergraduate degree was more humanities-oriented, then I pivoted to AI — I have an interdisciplinary background.
👦🏻 Koji
I remember you studied theology?
👦🏻 Yachen Song
I studied Hebrew and Arabic, and originally wanted to go in that direction.
👦🏻 Koji
Has your previous studies helped with your current work?
👦🏻 Yachen Song
Yes. Many of my decisions come from people I've met and information I've encountered.
Sixth, I am someone who likes reading, and also listening to audiobooks, though I'm not particularly into podcasts yet.
👦🏻 Koji
What's the most recent book you've read?
👦🏻 Yachen Song
Recently I read a very thin book, The Man Who Planted Trees, which Jinjian Zhang from Oasis Capital gave me — quite rewarding. I was also listening to Wang Dongyue's lectures on the Tao Te Ching a while back.
Seventh, I am someone who cares about "interesting." Whether recruiting, making friends, or dealing with investors, I value whether the other person has passion, whether they have moments where their eyes light up.
👦🏻 Koji
Who's the most interesting among your investors?
👦🏻 Yachen Song
This might offend some people (laughs). But most of our investors are quite young — the fact that they're willing to invest in us is itself a kind of "interesting."
👦🏻 Koji
What a diplomat.
👦🏻 Yachen Song
Eighth, I am someone who's not good at writing, including in both Chinese and English. Even though I spent eight years in the US, I've always struggled with written expression, so studying humanities was painful. For example, replying to WeChat messages, writing all-hands company letters, or drafting long emails to investors — I basically don't do any of it. But I'm good at chatting. When I worked on strategy at SenseTime, I often had to write PPTs, which was excruciating. I tried hard to overcome it, then realized it really wasn't my strength and gave up.
👩🏻 Ronghui
Then what do you do when you need to write prompts for AI tools?
👦🏻 Yachen Song
Also painful. I prefer interactive forms. In 3D space-time, in the future you shouldn't need to type — it should be "speak and it becomes reality." Like the floating little assistant in Genshin Impact, where you can just talk to generate content. That's a more natural interaction form, rather than a keyboard suddenly popping up for you to type. I'm really looking forward to 3D enabling this soon.
Ninth, I am someone who especially loves content creation. From elementary school I read all kinds of fantasy novels (Tang San, Wo Chi Xi Hong Shi, Cang Tian Bai He, Tian Can Tu Dou, etc.), then moved on to manga, anime, games — I even wrote fantasy novels and uploaded them to Qidian.
👦🏻 Koji
Can you be found on Qidian? What name?
👦🏻 Yachen Song
It was a tiny Qidian account with only 200 views.
But I did work on animation IPs at SenseTime, reaching millions of followers. I was also a KOL on Banciyuan. I like doing IPs and creating content.
👦🏻 Koji
How many followers did you have?
👦🏻 Yachen Song
Several tens of thousands at the time.
Tenth, I am someone who really wants to make games. In elementary school I had a very worn notebook where I designed several RPG games, with levels, equipment, inventory, exploration — you could even battle with classmates. After class many classmates would come find me to play, like an actual RPG.
The Origin of Everything: From Elementary School RPGs That Charged Fees, to a Dream of an Infinite World
👦🏻 Koji
Sounds like you had pretty high status at school.
👦🏻 Yachen Song
When you own a system, it's like you've created a small world — you become the "god" of that world with final interpretive authority. So people would come to top up, like giving me mushroom beef snacks, spicy strips, or those five-jiao dried tofu "Beijing roast duck" snacks, asking me to draw them stronger. It was fun.
👦🏻 Koji
Did that positive feedback from creating virtual worlds back then have a direct connection to what you're doing now with VAST?
👦🏻 Yachen Song
I think so. I especially love creating and making things — writing stories, building worlds, and also consuming other people's stories. That's my favorite thing, because the physical world has many limitations. The bigger world comes from human brains and imagination — that's an infinite world.
👦🏻 Koji
We're quite familiar with Simon. Beyond the previous 10, can you show us something we don't know? Please improvise another "I am..."
👦🏻 Yachen Song
At the same time, I'm also an older brother. I have a younger brother, and that's an important part of my identity. There's a big age gap between us, so I get to see more clearly what his generation likes. Ultimately, the products we build are for the next generation.
For him, AI is completely natural. When he runs into a problem, his instinct is to open GPT or DeepSeek. I still default to Google or Baidu. He's been consuming AI-generated content since he was a kid — it's just normal to him. When I was young, I was reading text. The internet wasn't mature yet; hardware and bandwidth couldn't support high-information-density media. Even though 3D is the most information-dense, most natural form, early internet had to rely on more abstract, lower-dimensional text.
So I started out reading novels on an MP3 player, hiding under my covers with a flashlight. The screen could only display ten characters at a time. To finish a five-million-word novel, I had to press the button half a million times. Then mobile internet arrived. In middle school, I bought an iPhone 4 and could finally consume images — the phone had a camera too. By college, we had more video content. It wasn't until I started working that short videos really took off.
👦🏻 Koji
Besides problem-solving — where you reach for Google and your brother reaches for ChatGPT — what other significant differences do you see between you two?
👦🏻 Yachen Song
I'm very used to reading WeChat official accounts; he's not. He prefers going to YouTube or Bilibili to find answers. For him, video is a higher-information-density format, which fits the environment he grew up in.
👦🏻 Koji
Have you already seen users who are "Born with 3D" in your user community?
👦🏻 Yachen Song
Yes, there's already a trend among younger users. My brother's junior, for instance, is a Bilibili creator making 3D content that I can barely understand — stuff like "Skibidi Toilet" and "Cameraman" — but he has hundreds of thousands of followers, and some videos hit millions of views. You don't even know who's watching, but clearly someone is. The generational divide is stark. The shift in information carriers from text to images to video to 3D, and in social forms from industrial society to digital society to intelligent society — you can't see it in a year or two, but over a decade, the change is massive.
👩🏻 Ronghui
Let's get specific about your 3D foundation model. Tripo 3.0 launched on August 20th, and you mentioned your focus on the next generation of users. Who is this designed for? And what's the biggest iteration compared to previous versions?
👦🏻 Yachen Song
We've actually been working on Tripo for quite a while now. It went live in early 2024, so it's been about a year and a half, and we've accumulated a lot of users. Currently we have three to four million professional creators globally, and over 40,000 enterprise clients, with more than 700 of those being large accounts. People are using AI for 3D, but previous products weren't truly pipeline-ready — they could only function in one part of the workflow, and still required professionals to modify and refine.
The epoch-making significance of Tripo 3.0 is that it's the first time we've reached a state where it can be directly used in most industries and scenarios. For example, you buy a 3D printer, put it at home for your kids, generate a 3D model on Tripo, drop it into the printer, and the print comes out great — no secondary operations or modifications needed. You don't even need to care about the model's structure or format, don't need to know how to fix 3D models, don't need to learn various DCC modeling software. None of that is the user's concern.
👦🏻 Koji
What work happened behind the scenes between 2.0 and 3.0?
👦🏻 Yachen Song
The workload was enormous, covering more data, better algorithms, and module optimization. This is a systematic engineering effort, not solving a single-point problem. Overall, there are significant improvements in controllability, success rate, fine-grained detail, and performance — with progress in geometric precision being especially critical.
👩🏻 Ronghui
Simon, could you give listeners something more concrete? I don't have a technical background, but before recording I discussed your materials with Gemini. Using Tripo 2.0 as an example, public information mentions you adopted a composite architecture fusing DiT and U-Net models. Gemini noted this was already quite challenging. I'm wondering — does 3.0 still use this architecture? If so, where does the challenge lie?
👦🏻 Yachen Song
3.0 doesn't rely on a single breakthrough point; it's multi-faceted system optimization. We developed a new representation called SparseFlex (SF). In April this year, we open-sourced Tripo SF, and the results are quite impressive. It significantly reduces the cost of generating 3D models and improves generation speed, because it skips the watertightness step while supporting generation in thousands of spatial dimensions with higher precision.
You can think of it as a 3D token representation. The better the representation, the higher the compression rate, reconstruction rate, and fidelity. This not only supports training on more data but also improves generation quality and precision. There have been similar representations recently — Mesh, NeRF, the recently popular Gaussians — and SparseFlex also performs outstandingly in AI 3D training.
👦🏻 Koji
Since the 3.0 launch, is there any data showing the value it's delivering to users? Usage volume, paid conversion, frequency of use?
👦🏻 Yachen Song
Both user volume and feedback have improved significantly. We've currently released the Standard version; an Ultra version will follow with better generation quality but longer generation times.

👦🏻 Koji
Are there application scenarios that 2.5 couldn't unlock but 3.0 can?
👦🏻 Yachen Song
You could put it that way. We have a product called Tripo Studio, which bundles a large number of AI algorithms with the goal of replacing traditional complex 3D production pipelines through an AI-native workflow. Since Tripo Studio launched on May 31st, revenue has more than doubled.
👦🏻 Koji
I remember you shared at a Crossing offline event, "AI Open Mic," that someone in Europe built a wrapper application on top of Tripo 3D's API and made quite a bit of money. Is Tripo Studio similar to this model?
👦🏻 Yachen Song
Not exactly. We want to make it more agent-like. Going forward, we'll add more dialogue boxes to the system, plus simple language interaction and drag-and-drop interaction. You can think of it as post-generation processing for 3D content: previously, when I generated something at "80 points," I'd typically do secondary editing in a traditional pipeline; now we complete that secondary editing on Tripo Studio, dramatically reducing the cost, barrier, and time of editing — that's the significance of Tripo Studio.
Tripo Studio introduces many globally unique new features. First is "universal automatic semantic segmentation." Previously, generated 3D models tended to be monolithic "wholes" that couldn't be edited after the fact — similar to tattoo images or tattoo videos, where after generation there's no source file, no way to edit in layers in Photoshop. Likewise, early 3D output was one solid block, difficult to separate into layers; now, generated 3D models can be automatically semantically segmented: the system understands the model's semantics, splits its components into independent blocks, and automatically completes and refines each component.
For example, if you generate a hand holding a bottle of water, the system can automatically separate the water from the hand, with each becoming a complete, independent 3D asset, stored in an asset library for replacement and reuse. This demonstrates Tripo Studio's capabilities. In other words, Tripo Studio has launched a complete suite of AI algorithms that defines the form and paradigm of AI-era 3D editing and interaction. Segmentation and part completion are very classic functions; competitors may implement similar features in the future as well.


👦🏻 Koji
You could understand this as similar to an important selling point when Canva launched AI — generated images could be edited in layers, adjusted layer by layer; Lovart also emphasizes this.
👦🏻 Yachen Song
This is indeed an innovation — no one had done it before, we achieved it first. We've also developed "universal automatic rigging." What this means: generated 3D models were originally static "sculptures," but now they can be automatically rigged and skinned. Beyond human models (which are relatively easier to handle), the system also supports automatic rigging and animation generation for cats, dogs, cows, snakes, fish, dragons, even octopuses and spiders. For example, after generating a dragon, the system can complete rigging at the claw or finger level, enabling it to move. This capability significantly saves creation time and lowers barriers.
Additionally, we've done low-poly generation. Traditional generated models often have hundreds of thousands or even millions of faces, which creates enormous local performance overhead in real-time rendering scenarios like games, XR, and the metaverse; when face count drops to hundreds or thousands, computational load decreases dramatically, enabling real-time rendering. For this, we developed our own autoregressive-based low-poly generation method, so output models naturally have lower face counts and greater practicality.
Similar features include Magic Brush smart brush and a series of other capabilities, which together constitute a complete workflow.

👦🏻 Koji
I see this as a strategic choice. Many teams might choose to go all-in on the base model, yet you spent significant R&D resources building Studio. Why?
👦🏻 Yachen Song
We believe that in the future, every large professional vertical community will have its own AI workstation — one that meets several criteria:
- End-to-end: it completes the full creative workflow in one place.
- High controllability: editable granularity must be fine enough to truly express creativity.
- Innovative interaction paradigms: no longer confined to traditional modes.
We realized this late last year and invested six months in R&D. On May 31, we launched the first version of Tripo Studio. It was still rough early on, but after several iterations, the results have improved significantly.
Model vs. Workstation: Why We Build Both the Engine and the F1 Car?
👦🏻 Koji
So you believe building only the base model isn't enough? Because some might choose to focus solely on the base model and leave something like Tripo Studio to ecosystem partners.
👦🏻 Yachen Song
Hmm, that's a great question. I believe the future will definitely involve both — base models and agents (or workstations, if you prefer). For example, Cursor will likely build its own base model.
👦🏻 Koji
So you see this as a defensively positioned agent?
👦🏻 Yachen Song
It's not defensive — it's a matter of logic. You can think of it this way: when building foundational large models, you're putting up a new wall; when building engineering and product features, you're patching the old wall. Because what you're solving are precisely the flaws that existed in the previous generation of models.
For instance, if the previous model didn't generate faces well, and I wanted to build an agent on top of it, I'd focus heavily on face optimization. But when the next-generation model releases and solves the face problem along with a hundred other issues, your previous effort may become obsolete.
So from a fundamental perspective, this is exactly the difference between AI 1.0 and AI 2.0. In the AI 1.0 era, the core was brilliant algorithmic scientists manually tuning parameters to train small models, then using these small models to solve various long-tail problems. For example, in the computer vision era, when we built smart city solutions, we'd have one small model dedicated to detecting whether garbage was dumped outside, another for detecting fights in prisons — very specific, granular problems.
In the AI 2.0 era, the core becomes data-driven. You train a general-purpose large model on massive data, hoping it can generalize to solve all routine problems.
Returning to the question — the advantage of building底层 models is solving most routine problems in one go. But why are people building tools, agents, workstations, or applications on top now? The core reason is: they believe AI 2.0 is dead, so they're still doing AI 1.0-era work.
Survival Rules for the AI 2.0 Era: Language Models Have Hit a Wall, But 3D Hasn't
👦🏻 Koji
You think AI 2.0 is already "dead"?
👦🏻 Yachen Song
It's not what I think — it's the problem those building them face. Think about it: if AI 2.0 were still rapidly evolving, with GPT-5, GPT-6, GPT-7 iterating one after another, Cursor would have no room to survive, because the problems it solves would all be covered by new models. Many agents are the same — they originally filled gaps in general-purpose models, but as large models upgrade, their value disappears.
👦🏻 Koji
So you believe models have hit a development bottleneck?
👦🏻 Yachen Song
Not all models — language models have hit a wall. And precisely because of this, many vertical applications and agents have emerged around language models, because the pace of language model development has relatively slowed.
But in 3D, the situation is completely different. You rarely see people building only applications, because without your own large model, the moment the next generation releases, your application becomes nearly obsolete immediately. It's like spending all that time patching the old wall, only for someone to put up a new wall and cover all your effort.
👦🏻 Koji
I understand that when large models are iterating rapidly, application companies "carving flowers" on top risk being drowned by the next upgrade. But I want to know — from your perspective, as a company that originally built base models, why are you also building applications at this moment?
👦🏻 Yachen Song
The core is this: we know better than anyone else where the next version of our model will iterate. We know which parts are worth patching on the old wall, and which parts don't need protecting because the new model will solve them.
This is our greatest advantage:
On one hand, by building tools or agents, we get closer to users, collect frontline feedback, and guide large model iteration;
On the other hand, we have the accumulation of large models and know the direction the model will take next. The combination of both creates a very positive cycle.
👦🏻 Koji
DeepSeek firmly refuses any commercialization — even when outsiders want to give them money, they reject it, because Wenfeng Liang believes this would dilute the team's focus on pushing the boundaries of model intelligence. So their ToC products are extremely simple; after going viral, they didn't maintain them, didn't even add server capacity. They put all their energy into base models.
Your choice is to do both base models and Tripo Studio, but your "granary" isn't as abundant as DeepSeek's. With limited resources, what considerations drive this strategic choice?
👦🏻 Yachen Song
I don't see this as attention dilution. On the contrary, if you only do base models, it becomes an academic form of "self-indulgence." Many things may be hot in papers but don't fully correspond to real user needs.
We don't "look for nails with a hammer in hand" — we start from real problems. This is also what makes our company special. I mentioned in my last sharing that we originally started with a 3D TikTok. But we hit a wall: we wanted to build a 3D UGC ecosystem and community, but discovered there was no 3D UGC in reality — only PGC. Why? Because there was no mass-market creator tool.
Just as without input methods, there'd be little text UGC; without smartphone cameras, little image and video UGC. In 3D, what's missing is a mass-market creation tool. So we set out to build AI 3D large models, aiming to lower the barrier and cost of creation. This original intention matters — we do this to solve real problems, staying with users and creators to see if our solutions actually work.
In other words, from day one, we weren't a "hammer looking for nails" company. Many large model companies start with technology, then find application scenarios; we saw user needs and pain points first, then decided to build large models.
👦🏻 Koji
So your real "nail" from the beginning was building a 3D UGC community.
👦🏻 Yachen Song
Exactly. Our goal is to build a mass-market creator tool that lets everyone create 3D content with zero barrier, zero cost, in real time.
👦🏻 Koji
So is Tripo Studio serving this audience now?
👦🏻 Yachen Song
Not yet, temporarily. You can understand Tripo Studio as currently serving professional users, oriented toward PUGC or PGC. We hope to gradually sacrifice some controllability and editing granularity, but in exchange for a large volume of content paradigms and templates. With these paradigms and templates, everyone can participate in creation.
For example, college students using Tripo Studio is fine, but for elementary school students, or my grandmother, it's still quite difficult. What we truly want to achieve is zero-barrier, zero-cost real-time 3D creation. Once we achieve that, there's opportunity for a UGC community to emerge.
👦🏻 Koji
What makes you believe the future is one where everyone makes 3D? Isn't this somewhat non-consensus? Because taking photos and videos feels natural, but 3D is higher-dimensional artistic creation — not something everyone would actively want to do.
👦🏻 Yachen Song
Actually, the "naturalness" you just mentioned has only existed for less than a decade. Video and photo capture becoming daily habits is only a recent ten-year phenomenon. Before short video emerged, how many movies did we watch per year? Before Xiaohongshu and Pinterest, how many times did we visit galleries per year? Before Weibo, Tieba, and Twitter, how many books did we read in a lifetime?
Returning to 3D — before 3D UGC content platforms existed, everyone was already gaming. Honor of Kings had 100 million DAU in its first year in 2015, and still maintains 100 million DAU ten years later — that's a truly mass-market product. Today's global gaming market is roughly $260 billion, already two to three times the combined size of publishing, galleries, and film markets.
Similarly, I believe future 3D UGC content platforms will be two to three times the combined size of Twitter, Weibo, Xiaohongshu, Douyin, Kuaishou, TikTok, Snapchat, and Instagram.
👦🏻 Koji
What content do you envision being consumed on this future 3D UGC platform? Games?
👦🏻 Yachen Song
That's a good question. When short video first emerged, if you asked people what they'd consume, there was no concept of short video yet — only film and video. So when we said we wanted to build a UGC video platform, people could generally only think of movies.
Similarly, now when people mention 3D interactive content platforms, they can only think of games. But future forms will definitely be far richer. Just as on Bilibili today, movies are only a tiny portion of all videos, and the short drama market has already surpassed film.
So I believe all the games we play today — whether it's Idle Fish King, match-three, Genshin Impact, Honor of Kings, or Game for Peace — will in the future be just a small branch under the larger category of "3D interactive content," an elite art form.

👩🏻 Ronghui
When you got to this point, I'd like you to give everyone a quick primer: compared to training a conventional large language model, what are the main challenges in training a 3D large model?
👦🏻 Yachen Song
We often use the "three elements of AI" to explain: if AI is like raising a horse, you need three conditions.
First, the horse needs grass — grass is data.
Second, the horse needs an excellent trainer — that's talent and algorithms.
Third, the horse needs a racetrack — that's compute. Only when you have data, compute, and outstanding algorithms (or scientists) can you train high-level AI.
In the 3D domain, these three elements differ significantly from language models. Starting with data. The early internet could only support low-information-density content, so the web accumulated massive amounts of crawlable text data. But the 3D internet remains immature, lacking publicly available large-scale datasets. This creates a natural bottleneck: where does the data come from?
We currently have the world's largest high-quality native 3D dataset, roughly over 40 million samples, approaching the level of the 3G-model "Monkey" in Black Myth: Wukong, as our training foundation. This is extremely important. By comparison, other major companies or competitors mostly remain at the million-sample level. We're the only team globally to reach the tens-of-millions scale.
👦🏻 Koji
Why is that? Are there some datasets that money can't buy?
👦🏻 Yachen Song
That's a good question — it falls under our core trade secrets (laughs).
First, on the "grass" level, we are indeed more abundant than others, which is crucial.
Second is "talent." Our team has fifty to sixty PhDs from Tsinghua University, all top-tier scientists. Assembling such a team isn't easy. Language models have been relatively prominent in recent years, attracting large numbers of NLP researchers, and OpenAI has long been deeply rooted in this field. But 3D is an entirely new interdisciplinary field, combining AI and computer graphics, and it lacks long-term accumulated expertise. Many researchers have only entered this field in the past year or two.
For example, two years ago at SIGGRAPH, the top graphics conference, you could still see many computer vision (CV) papers. But now at CVPR, the top computer vision conference, there are Best Papers strongly related to 3D. This shows 3D is gradually becoming a frontier direction, which also means early talent was severely insufficient.
So talent is especially critical. It's precisely these several dozen outstanding algorithm scientists that allow us to continuously iterate globally leading algorithms. This is an extremely rare and precious thing. Whether you can achieve this depends on entering the race early enough and committing decisively to the investment. Early entry means discovering and attracting talent sooner, cohering them into a team, and thus advancing rapidly. Data accumulation also relates to how early you started, though it's not entirely deterministic.
Third is the "racetrack" — compute and capital. We're currently among the most heavily funded companies in this赛道, with valuation also in the top tier. Having ample capital reserves, large-scale datasets, and top scientists theoretically enables producing excellent large models. But there's still an element of luck. As I mentioned earlier, finding the "oasis" isn't easier just because your caravan is larger.
👩🏻 Ronghui
Looking back at the past two years, how would you divide the key milestones or phases? For instance, at what point did you decide to build Tripo Studio and to do it quickly? Did this process also involve key people or key judgments?
👦🏻 Yachen Song
This can be understood as follows: our company has done only one thing in the past two years. Founded in 2023, our sole goal from 2023 to 2024 was to push our technology to the global frontier. We do have competitors, but observing them, many invested heavily in product early on. I believe the essence of early-stage product is technology, not surface-level features.
Here's an example: if your phone camera is 720p and mine is already 1080p, then adding various portrait, panorama, and infrared features on top of 720p is essentially meaningless. What you really should do is upgrade to 1080p or even 4K as quickly as possible, rather than obsessing over those add-on features. Adding features isn't the essence of product; technology is.
Therefore, for a long time in the past two years, our company didn't even have product managers — most of the code was written personally by the CTO.
👦🏻 Koji
But now you're building Tripo Studio. Does this imply you judge that 3D large models have hit a "1080p wall"? And at this point, you're diverting energy to applications rather than deepening the base model — aren't you worried that one day a competitor might release a "4K" base model?
👦🏻 Yachen Song
This has nothing to do with hitting a wall. We plan to advance in all directions simultaneously.
Tripo Studio was originally conceived as a UGC version first — a zero-barrier, zero-cost, real-time 3D content creation product for genuine UGC.
But why PGC first? The reason is that in early stages, both UGC and PGC have needs regarding generation quality (say, from 10 to 80 points). When models reach 90 points or higher, the needs of UGC and PGC diverge: UGC cares more about speed and getting things moving quickly, while PGC values fine details like topology and edge flow.
So the early goal was to ensure we could at least generate content that allows simple adjustments. Thus we first launched Studio for PGC, prioritizing existing users, then gradually covering native users — those who previously couldn't enter production pipelines but can now participate thanks to AI, such as AI-native users without 3D capabilities. Our strategy is to first serve existing PGC users, raise overall capabilities to the 80-90 point level, then consider how to serve UGC and PUGC (semi-professional users).
For this purpose, from late last year to this year, we built a product and engineering team of roughly twenty-plus people, specifically to solve product and engineering problems and support user adjustments and presets. For example, supporting model stylization (upload images to extract style), symmetry settings, T-pose and A-pose, and many other features.
We believe the primary task is to first serve our existing PGC users well, then gradually generalize to PUGC and UGC. In the generalization phase, beyond engineering and product, operations and growth become equally important. We haven't yet conducted large-scale operations and growth work, because existing professional users are highly focused on our product improvements, and product differences are obvious with relatively transparent information; thus users naturally adopt when effects improve.
But PUGC and UGC users won't continuously follow subtle improvements in large model performance. At that point, you need growth, BD (business development), or operations to communicate product value. Key scenarios for growth include: first, when information asymmetry exists and you need to proactively inform users; second, when products gradually homogenize and you need to build differentiation through brand and operations. Given that we still maintain differentiated advantages and professional users face no information gap, operations and growth will become focal points in the next stage.
👦🏻 Koji
Aren't you worried that building a complex product like Tripo Studio just to understand users is actually too laborious a path? After all, to achieve the goal of "understanding users better," you could conduct user research, or collaborate with application developers building on top of your base model.
👦🏻 Yachen Song
When you're building a foundational large model, you have almost no direct users. Your users are B2B enterprises; you must communicate with these B-end clients and rely on them to obtain user feedback.
👦🏻 Koji
But theoretically, you could also directly reach their users. Though perhaps not as straightforward, with some effort it's not impossible to get contact information.
👦🏻 Yachen Song
I'll say it again: the key difference lies in original intention. I don't know what DeepSeek's original intention is, nor do I know Moonshot AI's. Every company may have different goals. The reason we're doing this isn't for AI, 3D, or AGI per se, but stems from a clear original intention: we want to drive UGC for 3D content.
We've observed that creator communities, especially mass creators, lack a consumer-grade 3D creation tool, so we want to build such a platform. This is our most core starting point.
It's precisely because of this original intention that I'm doing this. You can't "have dumplings just because there's vinegar" — rather, it's the reverse: we set out to do this because the thing itself has value.
👦🏻 Koji
Indeed, a company's vision choice and strategic resolve along that path are both extremely important.

👦🏻 Koji
Sounds like Tripo Studio has performed well since launch? You mentioned it's already contributed over half of revenue. Earlier you said the goal of building Tripo Studio was to get more user feedback. So far, have you gained any new insights through it?
👦🏻 Yachen Song
We actually have a program called the "CEO Program" — Chief Experience Officer Program.
👩🏻 Ronghui
You've interviewed many users?
👦🏻 Yachen Song
Yes, roughly one to two thousand users have been interviewed so far, from various fields with very diverse use cases. What surprised me is that many real application scenarios I completely hadn't imagined before building the product. Initially we envisioned use directions as gaming, animation, and 3D content creation. But later we discovered many people use it for design, such as 3D printing and industrial design. So we gradually expanded our positioning to 3D content, 3D experiences, and 3D design.
Beyond that, there are large numbers of users from the art world, especially art school graduates. They use it for graduation projects, involving contemporary art, installation art, landscape art, new media art, and so on. Previously they lacked 3D creation capabilities; now they can achieve this through generative 3D tools.
There are also some users with disabilities who use Tripo Studio to express themselves and create. Then there's the XR application community — these users are very active, but in the past we focused more on XR hardware and less on software and application layers. Only through actual interviews did we discover there are many active XR developers globally, creating all kinds of interesting things every day: 3D presentations, 3D AI picture books, mini-games, and so on.
So we realized that 3D generation isn't just UGC — it may give rise to entirely new forms of play. In the gaming industry, there haven't been genuinely new gameplay forms for a long time; in the past decade, perhaps only Auto Chess counts as one. But 3D generation offers vast possibilities for new interactions and gameplay — something I hadn't previously considered.
The Ultimate Form of Technology Is a "Decompression"
👦🏻 Koji
Were these all discoveries you made after Tripo Studio launched?
👦🏻 Yachen Song
After launch, people started using it widely — some through the API, others through the SaaS interface. But essentially, what this reflects is that AI 3D has become a taken-for-granted capability.
People used to consider text-to-text, text-to-image, and text-to-video as natural and obvious. Now they feel the same way about text-to-3D. But if you think about it, going from "nothing" to "generating something out of thin air" is itself like magic.
The second point is how fast the technology iterates. Novel visual experiences keep emerging one after another, and people barely have time to process that this is actually a technology humans only invented two or three years ago, now already being deployed at scale in industry. Consider this: when the lightbulb and refrigerator were three years old, they were nowhere near mass adoption. Yet today, AI 3D already has millions of users and over 40,000 enterprises using it. That strikes me as remarkable.
More importantly, it expands the capability boundaries of ordinary people. Taking photos, shooting video, writing text — the masses could already do these things. But 3D modeling used to be possible only for trained professionals. Now everyone can "create something from nothing." That's no small thing.
Take the evolution of menus as an example:
- Before typewriters, all menus were handwritten.
- With typewriters came printed menus.
- With smartphone cameras, menus started including photos.
- Now you scan a QR code to order, or even interact directly on an iPad.
So why can't menus be 3D? If a restaurant could directly display the 50 dishes ordered by 10 people, using 3D models to show portion sizes and plating, you could immediately tell whether it would be enough to eat. But traditional 3D modeling was too costly to be practical. Today, if a restaurant could have a 3D ordering system for just a few dozen dollars a year, of course they'd be willing to pay.
Similarly, billboards and business cards could completely become 3D. The past forms of the internet were text, images, and video — but these were essentially dimensionality-reduced abstractions of the world, stopgap measures born of immature technology. When technology is sufficiently advanced, communication and expression will naturally return to 3D forms, the closest approximation to reality.
👦🏻 Koji
I think this paints a very imaginative future. The reason we watch short videos instead of 3D content today is simply due to constraints in technology, bandwidth, and device computing power — essentially, we're still "compressing" the world.
👦🏻 Yachen Song
Yes, "compressing the world" is an especially accurate way to put it.
👦🏻 Koji
Returning to my earlier question. You mentioned that you built Tripo Studio to gain user understanding and help the model iterate better. After several months of operation, what insights have you gained that you couldn't have obtained without Tripo Studio?
👦🏻 Yachen Song
There are actually quite a few. Internally, we maintain a requirements pool with hundreds of items, prioritized as P1, P2, P3, and so on.
A few examples:
- Users wanted to edit textures, so we developed smart brushes.
- Users wanted to modify geometry but didn't know where to start, so we explored whether natural language could directly edit geometry.
- Some wanted better topology in models, so we developed a retopology feature.
- Others pursued better hard surfaces and sharper corners, so we optimized our algorithms specifically for this.
- Some wanted to preserve brush strokes in textures and improve UV integrity, which we're also working on in dedicated R&D.
These granular needs have driven rapid iteration. On the other hand, major model updates have naturally resolved many issues — facial detail fidelity, hard surface performance, PBR support, texture quality, and so on.
So it still comes down to one principle: everything is in service of users and creators, not ourselves.
👩🏻 Ronghui
The world is essentially in a process of compression, and what you're doing is helping it "ascend in dimensionality." That sounds like a rather challenging path. Do you believe that as tools become simpler, people will actually choose to create in higher dimensions?
👦🏻 Yachen Song
I'd prefer to call it "decompression." The reason humans have constantly compressed is because they've been constrained by bandwidth, computing power, and other technical limitations. For example, early games could only use low-poly models because high-poly models with finer detail couldn't run, so everyone had to play games with rough graphics like Legend of Mir. But as technology developed, we could create Black Myth: Wukong, with polygon counts dozens or even hundreds of times higher — essentially a return to authenticity through "decompression."
The evolution of social platforms follows the same pattern: from Twitter and Weibo to Xiaohongshu, then to Douyin and TikTok — it's a gradual process of decompressing information expression. Humans originally expressed themselves in 3D, through statues and totems, before moving to murals, and then to text. Statues have the highest information density but aren't portable; text has lower density but spreads more easily — this was determined by technological conditions. The internet faces the same problem — a few bytes of data spread easily, a few gigabytes are harder, and 3D might take several terabytes, making it even more difficult.
So I believe the direction of technology isn't continued compression, but rather gradual decompression until we can finally work with the "source file" directly. And what is that source file? It's 3D. Video is essentially just taking one angle and one slice of time from a 3D world, but 3D is the complete source file. So I believe the internet will eventually reach an era where everyone can enjoy the source file.

👦🏻 Koji
Which competitor's news do you pay the most attention to?
👦🏻 Yachen Song
Hunyuan 3D has been doing quite well recently, and we're watching them.
👦🏻 Koji
Are there any competitors you particularly respect?
👦🏻 Yachen Song
We certainly respect all our competitors.
In the long run, the relationship may not be pure competition but rather "coopetition." Because the founding intentions differ: some teams aim to push the frontier of technology, others hope to influence industry through technology, and still others focus on tools themselves. For instance, certain well-known companies have goals that don't align with ours. To my knowledge, among all current competitors, no one is truly trying to do what we're doing. So ultimately, everyone will head down different paths.
👦🏻 Koji
What you want to build is a UGC-oriented 3D creator community.
👦🏻 Yachen Song
Yes. Our positioning is very clear: serving creators. Creators need community, so we serve them through community; they need platform, so we support them through platform. In other words, our goal is to build a complete network, not merely provide some single-point tool. You mentioned "decompressing all the way to the bottom" — that's a very apt description, and it's exactly what we're doing.
👦🏻 Koji
Aren't other companies also serving creators?
👦🏻 Yachen Song
Not quite the same. For example, some companies primarily serve game studios, providing customized solutions based on their needs. Others focus purely on tools. But we want to build a community and platform — a completely different path. The difficulty of this is immense, and the investment requires long-term persistence; it can't be completed in the short term.
👩🏻 Ronghui
Then how do you view commercialization? Serving game studios sounds like a model closer to revenue.
👦🏻 Yachen Song
I don't actually believe serving game studios is closer to commercialization. Looking globally, companies that have long focused on serving game studios haven't fared particularly well. Try thinking of counterexamples: is there any company that became exceptionally successful by focusing on serving game studios? They're almost impossible to find. In other words, the B2B path isn't the most ideal in this industry.
👩🏻 Ronghui
So what does your own commercialization path look like? Over the past two years, has it proceeded according to plan, or have there been surprises?
👦🏻 Yachen Song
I believe that in the early stage, the essence of commercialization isn't "commercialization" itself, but product. As discussed earlier, the essence of early product isn't features, but technology. Only when a product truly possesses sufficient differentiation does commercialization have a solid foundation.
Specifically, growth and BD are most valuable in two situations: first, when products are highly homogenized; second, when there's massive information asymmetry among users. If neither condition holds, then the core task is to refine the product further. For us, what's more important now isn't hiring more BD staff, taking clients to dinner, or buying ads for growth — it's making the product itself stronger. Discussing commercialization methods becomes more appropriate at the next stage.
👦🏻 Koji
So can I understand this as: you feel the product isn't good enough yet to warrant large-scale promotion?
👦🏻 Yachen Song
It's not that "the product isn't good enough, so we can't promote it," but rather that "making the product better is more critical than promotion." Take NVIDIA, for example — their products are extremely powerful, but they don't rely on sales teams to drive commercialization, nor do they need to constantly pitch ByteDance or Baidu. For them, continuously improving the product is more sensible than additional promotion. We're the same: when customers have no significant information asymmetry, and the product is sufficiently differentiated, the optimal solution is to keep strengthening the product rather than prioritizing BD or buying traffic.
👩🏻 Ronghui
I've noticed in several interviews that you mention sharing the company vision with the team at year-end. You said the content doesn't actually change much year to year. Could you elaborate on what has remained unchanged over your two-plus years of entrepreneurship, and what has changed?
👦🏻 Yachen Song
We do an all-hands share around mid-March each year. I usually just need to screenshot last year's slides, drop them into the new version, make minor updates, and keep presenting. That's because the vision and core path have barely changed — we've simply advanced another year and accomplished more goals.
👦🏻 Koji
So compared to your angel-round pitch deck, not much has changed?
👦🏻 Yachen Song
Strictly speaking, we didn't have a formal pitch deck at the angel round stage, but the internal presentation slides we used then aren't dramatically different from what we have now.
👦🏻 Koji
Very interesting. I recently heard a story: Huiwen Wang met with an entrepreneur and asked him to pull out his angel-round pitch deck from four or five years ago, then compare it to today — see what changed and what didn't. It turned out quite a lot had changed after pivoting.
👦🏻 Yachen Song
We're very simple on this front because almost nothing has changed. Whenever I want to adjust the vision, the team tends to push back. So the vision and path have remained stable. Team members have strong understanding and identification with the vision and path — this is the source of our cohesion. We have virtually no talent attrition, and the reason is that everyone has deep confidence in and passion for this long-term path. In half a year, I'll do my fourth share, and I hope I'll still be presenting the same content.
👦🏻 Koji
Where does this sustained sense of identification come from? Is it because everyone is themselves a potential user of the product, looking forward to using something like this? Or is it for other reasons?
👦🏻 Yachen Song
Some of our colleagues are indeed potential users. Others have stayed because they've witnessed the vision materialize step by step. Many things have progressed even faster than we initially imagined. There's something inherently powerful about watching a story become reality. The vision we laid out five years ago now sees new validation every year, which deeply energizes the team and further solidifies everyone's conviction.
Welcome to the Fourth Sector: When Experience Becomes the Sole Measure of Value
👦🏻 Koji
Can you describe what the world would look like if your vision were fully realized?
👦🏻 Yachen Song
We typically say human society has moved through the agricultural age, the industrial age, and into services — what we call the three sectors. But I believe there's actually a "fourth sector" — the content sector. Its core is the creation of content and experience.
In the future, the value humans can create in the physical world will become increasingly limited. Because virtually all physical value will be produced by robots. In other words, it will be hard to find things that "humans can do but robots cannot." If so, where does human value come from? I believe the answer lies in creativity.
How do we measure the value of creativity? There's one metric that defines it: the total time all people spend in the content and experiences I create — that is the value I create in this world.
👦🏻 Koji
That sounds a lot like those games you designed in school as a kid, drawing classmates to come play with you during breaks.
👦🏻 Yachen Song
The "content and experiences" here don't have to be games — they can take any form. Maybe it's battling or reveling together, maybe it's meditation and exploration. I can't predict what people will specifically choose in the future, but what I can be certain of is that human nature pursues "optimal experience." When people have enough choices — even infinite choices — they can exercise the freedom of "voting with their feet" to find their own optimal experience.
So how do we achieve infinite experiences? The answer is UGC and AIGC tools. UGC provides a constant stream of creation, while AIGC tools multiply creative efficiency. Take Douyin as an example: short video content approaches infinity, and with recommendation algorithms on top, the user experience naturally improves. This mechanism is also fundamentally fair — people vote with "time and attention," determining which content holds more value.
Therefore, in the future, the wealthiest people may not be those who control the most land or capital, but rather those with the most creativity. They can create a world where everyone wants to stay. It might even be something as seemingly simple as a "thief simulator" that, because it offers a unique experience, draws countless people to immerse themselves in it.
The core of this value remains "experience." If recommendation algorithms are better, they can more precisely match users with content; if computing power is stronger, it can deliver a smoother experience. There's a scene in the series Upload where in the virtual world, those who pay less experience lag, while those who pay more enjoy ultimate fluidity. The future world will likely follow similar logic.
When computing power, recommendation algorithms, and content creation tools (like Tripo, Midjourney, etc.) come together, people can continuously generate new content and immerse themselves within it. Those with stronger creativity, better tools, and more computing power will create experiences that truly bring joy to the masses. And this, too, will become a new source of wealth and value in the future.
👩🏻 Ronghui
What's your personal experience of founding an AI startup these past two years?
👦🏻 Yachen Song
I feel incredibly fortunate. We had enough resources to cold-start, but not so much that we'd be trapped by capital's "resource curse." No great company in history was born during a period of extreme capital market exuberance. Today's environment happens to maintain a healthy sense of hunger — something to be grateful for.
👩🏻 Ronghui
If you could say one thing to yourself from two years ago, right when you decided to start this company, what would you say?
👦🏻 Yachen Song
I'd say: "You're amazing, you really have guts. Starting a company was the right call, you chose correctly."
👦🏻 Koji
So have you actually suffered through this entrepreneurial process?
👦🏻 Yachen Song
Of course. There are problems to solve every single day when you're building a company — that's inevitable.
👩🏻 Ronghui
How do you work through that pain?
👦🏻 Yachen Song
I play games, like Dungeons & Dragons.
When the physical world is your entire world, pain occupies your whole life. But if the physical world is only 50%, with the other half filled by the virtual world, then pain's proportion decreases.
If you care so much about the world God created, then the world God created will dictate all your moods. But when you also have worlds created by people, then the world God created is just one part of your life.
👩🏻 Ronghui
I remember you mentioned in another interview that you're someone who really needs the virtual world.
👦🏻 Yachen Song
Yes, absolutely. I think everyone actually needs it — it's just that many people haven't realized it yet.
👩🏻 Ronghui
Thank you Simon for sharing your thoughts on 3D foundation models, tools, and entrepreneurship with us — truly wonderful.
👦🏻 Koji
Thank you, looking forward to having you back next time.
👦🏻 Yachen Song
Alright, bye.
🚥
