2025 AI Scene: What We Witnessed and Imagined This Year
There is no turning back — the only way forward is through.
There's no turning back now. We can only keep moving forward.

👦🏻 Podcast interview: Koji, Ronghui
🥷 Edited by: Starry
🧑🎨 Layout: NCon

This week on Crossing, we're joined by Minghao Zhuang (host of Tulong Zhishu) to review the major events of 2025 in AI and tech, and to share some of our own memories and feelings from living through it all.
This year, we've been witnesses — watching technology iterate at breakneck speed, watching products upend daily life in unprecedented ways. At the same time, we've been swept up in a collective fever dream about the future, oscillating between exhilaration and bewilderment in the face of innovation's tidal wave and endless possibility.
We start with Minghao's keyword for the year: "inflection point." From DeepSeek R1 at the start of the year to the recent Sora 2, we revisit the model wars, Manus igniting the "Year of the Agent," open-source ecosystems and talent flows, and finally turn to the capital markets — how should we understand this collective fever dream about the future?
🎧 Listen on WeChat:
🎧 Listen on Xiaoyuzhou:

🎥 Watch the video podcast. Recorded at the Shanghai AI Hacker House, also available on WeChat Channels, Xiaohongshu, Bilibili, and YouTube.
Full transcript is lengthy (19,866 characters). Here's the table of contents:
🟢 2025: An Inflection Point, Up or Down?
"We've unwittingly walked into the limits of many things — technology, products, money."
- Trying to sum up 2025 in one word: why "inflection point" fits best
- The upward inflection: the data center construction frenzy foreshadows a 2026 explosion
- The downward inflection: when does the bubble burst? Have we already hit the limits of technology and growth without realizing it?
🟢 The LLM Battlefield: Divergence and Consensus in China-US Tech
- How did DeepSeek R1, built for millions, upend the hundred-billion-dollar infrastructure narrative?
- Sam Altman redefines the "Turing moment": why AGI might not arrive with a bang, but as a quiet step forward?
- Survival rules for top model makers: Anthropic goes deep on B2B, xAI takes the wild path, why was Microsoft forced to build its own models?
- Is the pure chatbot war already over? Behind ChatGPT's 800 million weekly active users — moat or growth ceiling?
- Chinese players' consensus and weapon: why did "open source" become the only means to counter the US AI trend?
- What does DeepSeek V3.2's release mean? Why we might not see V4 or R2 this year?
🟢 The Other Path to AGI: From Sora 2 to World Models
"If vision models are also at the main table, they might achieve AGI through a completely different route."
- Why is the multimodal battlefield more competitive than language models? Meitu, marketing video agents... use cases and monetization paths are crystal clear.
- OpenAI's product philosophy: why did Sora 2 reach millions of households, not other equally advanced products?
- Does the world really not need an "AI Douyin"? Perhaps OpenAI never aimed to build one in the first place.
- Google is back! Could the world model Genie be another path to AGI, even "the womb of worlds"?
🟢 The Year of the Agent, Then What?
- Why will Agent linger at the L3 stage for a long time? For the first time, it extends AI capability from "language" to "action."
- Manus's historical significance: it showed users for the first time what an Agent should look like — defining the mental model is worth its weight in gold.
- How do Agent startups survive? When general-purpose Agents become the giants' gospel, vertical domains like law, finance, and marketing thrive unexpectedly.
- The attention economy's squeeze effect: when mainstream tracks get crowded, why can even niche tracks like "AI dynamic comics" — with limited ceilings — still raise money?
- Why couldn't Siri become the true "phone assistant," but on-device Agents today can?
- The HarmonyOS HMAF framework's revelation: when the OS stops doing everything itself and delegates intent to apps' native Agents, what new opportunities emerge for developers?
🟢 Open Source: An Ecosystem with Chinese Characteristics
- From "top models must be closed" to "open-source models will peak in 2026" — why did Sam Altman change his tune?
- How does open source monetize? DeepSeek's API still sells, at a fraction of OpenAI's cost.
- How does open source become a "weapon"? In competing for Europe, Southeast Asia, the Middle East and other middle grounds, open source naturally carries trust advantages.
- How strong is local deployment demand? A laptop with massive RAM and VRAM sold out instantly because it was ideal for running LLMs locally.
- What new "ecological niches" can developers seize? HarmonyOS developers made 70,000 yuan monthly from several small apps.
🟢 Secondary Market Frenzy, What About Primary Markets?
"Back then people thought AI was a technology, an industry. Today, AI is the market itself."
- Sam Altman's "endgame thinking": what happens when a company tries to package the next five years of growth expectations?
- On the flip side, Chinese VC's familiar mobile internet growth narrative "cannot be replicated today."
- How do investors find conviction? When the pure AI software story stops working, everyone pivots collectively to hardware founders with DJI, Roborock, and Dreame pedigrees.
- The ultimate question versus the dot-com bubble: fiber optics could pave the future, but GPUs that become obsolete in three years?
- An interesting signal: besides NVIDIA, the best performers in the S&P 500 this year were Seagate and Western Digital — hard drive companies.
- Bubble alert: when AI giants start relying on debt financing, is the ghost of the "subprime crisis" drawing near?

2025: An Inflection Point, Up or Down?
👦🏻 Koji
This week on Crossing, our guest is old friend Minghao Zhuang. Today we'll review the major tech events of 2025 with him, share some AI-era memories, and talk about how it feels to be living through it all.
Honestly, while preparing this episode, my feeling could be summed up in one word — fast. Whether it's technology iteration, product updates, or shifts in the global landscape, 2025 feels like it's running at 8x speed. Last year people often said "one year in AI equals ten years in the human world." Fewer people say it now, but the feeling persists.
Minghao has been documenting this era, constantly making presentations — I've given him a nickname: "the Sima Qian of the AI era." Our first question for this "Sima Qian": how do you think future generations will talk about this year? Will it be the year of the bubble, the year of inflection, or some other keyword?
👦🏻 Minghao Zhuang
One word comes to mind that fits perfectly — inflection point. The beauty of an inflection point is it can bend upward or downward.
If it bends upward, take the data center construction everyone's been discussing. By all projections, 2026 is when things explode — that's the upward possibility. But there's another conversation: will the bubble burst? If we enter a bursting phase, we might start heading downward by late 2025. Come year-end, you'll find we've unwittingly walked into the limits of many things — technology, products, capital, and larger dimensions beyond. AI has reached a kind of human limit.
The LLM Battlefield: Divergence and Consensus in China-US Tech
👦🏻 Koji
We'll review 2025 in several sections today. Too much happened this year to cover everything, so we'll inevitably filter through our own experiences. Some events hit us harder personally, but others might think, "how could you skip something this big?" That's unavoidable.
Let's start with LLMs, multimodal, and Agents. A massive early-year event was DeepSeek's R1 release. It was a global sensation, wall-to-wall coverage. Even today, I feel it profoundly changed how we live and work.
Looking back, Minghao, what's your take on R1? Where does its significance lie?
👦🏻 Minghao Zhuang
R1 officially launched around January 20-something, two days before the Lunar New Year holiday. At that moment, the most watched person in China was Yocar — founder of Black Myth: Wukong. He posted on Weibo calling it potentially "a national-destiny-level milestone," sparking another round of discussion. For those few days, virtually every major US financial and tech outlet was covering DeepSeek.
Meanwhile, the US had already been building its narrative around the next wave of AI infrastructure investment, with headlines constantly throwing around "billions, tens of billions." DeepSeek R1's paper mentioned its final training run cost roughly single-digit millions. That stark contrast generated enormous debate.
From an investment perspective, the US "Magnificent Seven" and NVIDIA's stock both took major hits, because people began questioning the established narrative — maybe there was another way to solve this. Of course, new trends emerged afterward. Like the coal industry back in the day: more resource extraction created more opportunities. Short-term volatility, but long-term growth continued. That narrative persists to this day.
Starting with DeepSeek R1, we saw the competition between China and the US in the large model space — spanning technical approaches, the open-source versus closed-source debate, and the scale of product deployment.
The industry commonly uses an L1-to-L5 implementation path: L1 is the familiar chatbot, L2 is reasoning. The O1 model launched in September 2024, when all the major players were working to replicate either R1 or O1. O1 was the representative reasoning model, and within two or three months of R1's release — roughly Q1 2025 — virtually every leading vendor had introduced their own reasoning model. To this day, everyone continues updating their base models and reasoning models.
So if we look back from the end of 2025, R1 essentially defined the competitive trajectory for the year. What followed was mostly engineering optimization and refinement.
👩🏻 Ronghui
What other model updates are you tracking?
👦🏻 Minghao Zhuang
The other breakout moment that cycle was of course GPT-4o's image generation, though it also triggered copyright issues and other problems — very similar to what Sora 2 is facing today. If you look only at OpenAI's technical roadmap, in 2024 people had sky-high expectations for GPT-5, believing it would be the milestone that crossed the AGI threshold. But when GPT-5 arrived this year, it didn't meet those expectations.
I was listening to Sam Altman on the a16z podcast the other day, and he said a lot of the changes we're experiencing may not be as dramatic as before. Even when GPT-3.5 powered ChatGPT, it arguably crossed the Turing test in some sense, but reality didn't get "turned upside down." Two-plus years later, we just stepped over it gently and kept moving.
Right now it's also hard to define what AGI even is. From November 2022 when ChatGPT launched to now — two and a half, almost three years — pessimistically speaking, progress has remained largely on a linear track overall. This year, various vendors have continued updating their models. Google has moved extremely fast — whether the Gemini series or visual models like Veo 3, they've been very strong.
👦🏻 Koji
We actually just published a WeChat article, on why Google is "back".
👦🏻 Minghao Zhuang
Right, and Claude has also been steadily doing its own thing, with even more rhythm in 2025. From a market economy perspective, it's the number two among creative companies, and number two has to differentiate from number one. While the underlying model gap isn't large, Claude's scene selection has become increasingly clear, focusing heavily on B2B, coding, and similar scenarios. Its growth curve is actually faster than OpenAI's, though from a smaller base.
This year xAI's Grok series updates have also had distinctive features, like the partnership with X (formerly Twitter) and virtual companion applications.
👦🏻 Koji
Including now after Sora 2 came out, X also quickly launched their own video model, and the approach is pretty "wild." Because it can basically generate photos of anyone, with zero respect for privacy.
👦🏻 Minghao Zhuang
And you can even write "spicy" in the prompt, asking for something "hotter."
Right. Behind this it's not just a technology issue, but also a business issue. Microsoft released its own model at the end of August. Due to its complex relationship with OpenAI, Microsoft had to build its own model. That's roughly the landscape among the core US players — OpenAI, Google, Anthropic, xAI, and Meta. Meta hasn't made much noise this year, mainly just hiring like crazy.
👦🏻 Koji
Meta's only headlines are "writing big checks to hire people."
👦🏻 Minghao Zhuang
What about China? The original "Six Little Dragons" dynamic has essentially ended. The current consensus is: open source.
Building an open-source ecosystem may be China's only means of finding a path distinct from the US approach to AI. The two top open-source players are DeepSeek and Qwen, running almost neck and neck with rapid growth.
This year, DeepSeek V3.2 was an important marker. It means we probably won't see V4 this year. DeepSeek V3 was the base model released at the end of last year, which through reinforcement learning became R1. After V3.1's continuous updates, people assumed the next would be V4, but instead we got V3.2. So the speculation became that R2 may not arrive this year either.
Qwen's story, meanwhile, is tightly connected to Alibaba's overall AI strategy, covering development investment, language models, multimodal capabilities, and coding. Additionally, Moonshot AI and Zhipu AI have both adjusted their strategies, releasing Kimi K2 and GLM 4.5/4.6 respectively, with new explorations in coding, agents, and open source.
In Europe there's also a company called Mistral that's performed notably in natural language models. These constitute the current mainstream large model landscape. Of course, we haven't yet delved into multimodal, coding, and agents here.
👩🏻 Ronghui
Looking back, what do you think will be remembered ten years from now about this year?
👦🏻 Minghao Zhuang
I think R1 will definitely be remembered, the inflection point of pushing models forward will be remembered, and the representative of agents will be remembered. On the multimodal side, if I had to pick just one, I believe it would be Google's Veo 3 — the first model to go from silent video to video with sound, which is a huge leap.
Looking further ahead, world models are also accelerating this year, with more and more players getting involved. It's still in a relatively demo-heavy stage, not yet at a point worth deeper discussion. But if forced to predict, there may be an important inflection point around late 2025 or early 2026.
👦🏻 Koji
Right, actually as Sam mentioned, recent releases have mostly been incremental changes. Many updates, especially to large language models, no longer excite people as much. Against this backdrop, if another "world's best" model appeared right now, Minghao, how would you view it?
👦🏻 Minghao Zhuang
First, I'd look at which leaderboard it took "world's best" on. There are several different evaluation systems now: blind tests, scoring, benchmark testing, and so on. Results vary under different standards. So indeed, for both the general public and industry practitioners, landing a so-called first place or SOTA has very limited shock value at this point. Such news is mostly "managing up."
This year, voice model competition has also been exceptionally fierce — we're in the voice business ourselves, so we can feel the intensity. Voice was originally a relatively marginal battlefield, but its advantage is the lower investment required. For vendors that have just entered the arena and want to build reputation or influence, voice is a cost-effective choice. Beyond voice scenarios themselves becoming more numerous, cost-effectiveness is the key consideration. So looking back at the good PR pieces this year, voice vendors are often not the biggest players, but they still treat AI as a significant strategy.
They typically have something like an AI Lab internally, established relatively late, but the company is willing to invest. To justify the department's existence, they must find a breakthrough in the short term, and voice appears to be exactly that opportunity.
👩🏻 Ronghui
The window for technical leadership is getting shorter and shorter. Everyone says AI vendors didn't even get a good National Day holiday this year. In your view, where does users' willingness to pay for chatbots mainly come from now?
👦🏻 Minghao Zhuang
It depends on usage habits, willingness to pay, and brand perception. This year there's been constant emphasis on improvements to context memory length — can these accumulations create a flywheel effect? I think it's already starting to work.
Returning to the pure chatbot battlefield, this battle is essentially over — ChatGPT won. Even though Google Gemini has grown nicely over the past two quarters, it's only gone from a fraction of a percent to a few percent. ChatGPT holds the vast majority of users; its latest reported weekly active users is 800 million. Looking at the growth curve Sam published, the increase is staggering. About a year ago, that number was only a few hundred million — it doubled in less than a year.
Earlier this year, a journalist asked Sam: would you rather have the most cutting-edge AGI technology, or a platform with 1 billion users? He answered "both," but clearly leaned toward the latter. We can see that many things are hitting ceilings. User volume has reached a ceiling, fundraising and other links have also hit bottlenecks. In the pure chatbot arena, ChatGPT has built strong brand recognition, and continues strengthening memory and comprehension capabilities.
I read a great article: What OpenAI is doing is the "grand unification" that Meta and Google failed to achieve, but WeChat accomplished. In China, we call this the "All in One" strategy. At this OpenAI launch event, when Sam introduced ChatGPT, he mentioned it can directly invoke services like Spotify, Zillow (real estate), Canva (design), and Figma (UI design) within the chat interface — this is essentially a "mini-program" ecosystem.
From a product perspective, OpenAI has not given up on the possibility of becoming an "All in One" platform.
👦🏻 Koji
I think this ultimately comes down to the different backgrounds of Sam Altman and Anthropic's Dario Amodei. It manifests in strategic choices and core competencies — OpenAI remains a company with very strong product sense and strategic capability.
Recently everyone's been discussing Sora 2, but many have already forgotten ChatGPT Pulse, released just four weeks ago. Pulse is an important product for boosting user stickiness — it enables better utilization of memory features, helping convert users from weekly active (WAU) to daily active (DAU). From a product implementation standpoint, Pulse performs very naturally and smoothly.
This also explains why in the secondary market right now, "wherever OpenAI points, stock prices rise." Behind this lies a question: why OpenAI, and not Google or Anthropic? This still reflects Sam Altman's strategic judgment and ability to shape narrative. He's very skilled at converting these abilities into company advantages.
Of course, this "all rise together, all fall together" model also means the bubble may eventually burst. But if that day ever comes, OpenAI will likely be the last one standing.
👦🏻 Minghao Zhuang
Many people forget that Sam was originally an investor. He mentioned on an a16z podcast not long ago that he's not great at management, but better at investing. Inside OpenAI, he actually operates more like an investor: nurturing, incubating, and supporting promising teams to help them grow. That's very similar to early-stage investing.
Sam was a partner at Y Combinator (YC), not a traditional VC at a dollar-denominated fund like Benchmark or Sequoia. He was a partner at YC. Before running OpenAI, he was best known for his startup course at YC. I studied technology economics and management, also called entrepreneurship management, in college. I once suggested to my advisor that graduate students' first class should just be Sam's YC lecture videos — that was already the pinnacle of startup education. It's just that many people have forgotten this now.

👦🏻 Koji
Let's start with DeepSeek R1 and look back at it now. There was an assessment at the time that said "Chinese people are good at getting big results with small money." When R1 launched, Sam Altman had just signed a massive contract — the "Stargate Project" — and was mocked by quite a few people. The contract was worth $500 billion.
Everyone was saying: "Look, DeepSeek pulled it off. Americans only know how to burn money." This was also an important reason for OpenAI's stock price drop at the time. But looking back, several months or even half a year later, do you still think the "get big results with small money" logic holds? The United States still seems to be spending wildly, and increasingly so.
👦🏻 Minghao Zhuang
Although I'm not a technical person, looking at developments over the past few years, large model R&D has increasingly become an engineering problem. The key lies in trade-offs and strategic choices, not zero-to-one breakthroughs. In this situation, the American approach is "spend big to stack resources," while China more often advances through "overtaking on the curve" or "clever workarounds." But the two approaches are gradually converging. When we find new methods, they learn from them; when they achieve breakthroughs by throwing money at the problem, we try that too. As attempts multiply, costs naturally rise.
Over the past year, while DeepSeek hasn't expressed much externally, if you look at Qwen's logic, it's already closer to the American approach. Eddie Wu talks more about CAPEX (Capital Expenditure) investment. Although Qwen has some "clever tricks" too, the overall narrative blends both logics. Who's right and who's wrong? There's no absolute answer — they learn from each other, imitate each other, evolve.
But one thing is clear: among algorithms, data, and compute power, only compute can be solved by throwing money at it. Algorithms are especially difficult, data somewhat less so.
👦🏻 Koji
And talent — that can be attracted with money too.
👦🏻 Minghao Zhuang
Some say no company has ever gone bankrupt from paying genius employees high salaries. When we unconsciously reach limits in many things, we may discover there are other ways to achieve the same results with less money.
Mark Zuckerberg has recognized this. Meta invests tens of billions of dollars annually, and this "arms race" may continue for several more years. Although they have abundant cash, even if they poached all the talent in Silicon Valley, they couldn't win just by spending money.
So his strategic shift has its reasons. For example, Mira Murati's co-founder at Thinking Machines went directly to Meta. That company is already valued at $12 billion, and he holds at least 10% — worth over a billion dollars. And Sam's offer was $3.5 billion, obviously higher.
👦🏻 Koji
That's insane. We'll talk more about China-US competition later. I do think the United States is investing more money in this area now.
👦🏻 Minghao Zhuang
In the conversation between Guangmi and Xiaojun last quarter, it was put very clearly: American financial capital plus Jewish financial wisdom plus Chinese engineers — that's the biggest representation of this AI wave. When you bind these two together, you can explain all phenomena in the current AI industry. At least in the United States, everything is obvious.
Another Path to AGI: From Sora 2 to World Models
👦🏻 Koji
We just talked about large language models, and now I want to discuss the multimodal battlefield. This year Google first released Veo 3, then OpenAI launched Sora 2, achieving exciting breakthroughs in both model progress and productization. At the same time, virtually every major player is investing in this space.
What I'd like to discuss is: in such a red ocean of fierce, homogenized competition, what are the decisive factors for victory?
👦🏻 Minghao Zhuang
When we talk about multimodal now, images and video really can't be separated. And on this battlefield, China-US competition is even more intense than in pure language models. Chinese manufacturers have shown very strong capabilities at this stage. Platforms like Douyin and Kuaishou, as well as some startups like Vidu and PixVerse, have all formed a head-to-head posture in this segment. For example, when Veo 3 just launched its audio-visual synchronization feature, within less than three months, Keling AI and ByteDance's models caught up. As long as the direction is clearly defined — whether product boundaries or technical routes — both China and the US will rapidly follow up.
When Sora 2 launched, people immediately started discussing when a domestic equivalent would appear, and you could see prototypes in about two months. Because the application scenarios in this segment are very clear, it penetrates industry chains immediately. Language models still need to face complex industry validation in law, finance, HR, and other fields, while image and video applications are more direct and commercially promising. Just look at Meitu's stock price over the past year to feel this trend. Now many companies are building AI Agents for video marketing — these scenarios are already very mature and no longer need additional justification. Once model controllability or costs reach a certain threshold, penetration and diffusion happen immediately.
Overall, this has already become a systematic competition. China and the US also differ in resource endowments. China's short video ecosystem is extremely rich, with tighter integration across e-commerce, marketing, and toolification. This iteration of image and video toolification also continues the logic of the mobile internet era. Think back to when the App Store launched — photography apps have always been the most competitive category, with photography-related apps selected as Best of the Year for multiple consecutive years.
👦🏻 Koji
Right, and there are always new apps emerging.
👦🏻 Minghao Zhuang
Yes, the battle continues to this day. And precisely because demand in this track is so rigid and scenarios so numerous, the real challenge for manufacturers is "what exactly to build." Technical evolution itself is certainly difficult, but the more critical question is how to land it. B2B or B2C?
After Sora 2 appeared, some people saw it as "AI Douyin," but others interpreted it positively — it's actually using interaction patterns that the mass public already finds familiar, making AI enhancement feel like a natural experience. OpenAI packaged a clearly bounded product through a seemingly simple implementation path in a short time — that's actually very difficult. Adding features is easy, but making trade-offs, polishing, and packaging within appropriate boundaries places extremely high demands on product managers' insight, user understanding, and interaction design capabilities.
👦🏻 Koji
Indeed, I really admire OpenAI. Actually, they're not the only company with this level of model capability. For example, we haven't mentioned MiniMax yet — they launched "Hailuo AI" video model half a year ago, and had a very viral hit: kittens competing in Olympic diving, with gymnastics movements reproduced at world-leading fidelity. But after so many manufacturers produced top-tier video models, only Sora 2 truly brought AI models into every household. It made AI completely break out of the inner circle.
👦🏻 Minghao Zhuang
Yes, though some would argue that's partly because it came from OpenAI. Actually some third-party manufacturers had attempted similar things earlier, perhaps more in images than video. But because they weren't OpenAI, weren't a top-tier company, they didn't get equivalent attention.
Previously in the multimodal space, one of the "Six Little Dragons," StepFun, besides working on model projects, also developed an image community called "Lipu." The team later disbanded, but early data performance was actually decent — retention, activity levels weren't bad. It just couldn't continue for various reasons. This doesn't mean no one was trying at the time; it's just that people's optimism level wasn't high enough then. Now after Sora 2 appeared, that optimism value may have risen from 30 to 60 points.
👦🏻 Koji
Right, at that time projects like "Nie TA" and "Lipu" were still stuck in the ACG subculture circle, without truly generalizing. On one hand people couldn't quite understand them, on the other hand they felt the ceiling was limited.
👦🏻 Minghao Zhuang
Exactly. So when Sora 2 appeared, my assessment was actually quite simple. Many media headlines asked "Do we really need an AI Douyin?" I think that's a fair question. We already have a mature Douyin ecosystem — from content, ecosystem, retention, interaction to monetization, it's already incredibly complete. The world truly doesn't need another "AI version of Douyin."
But for OpenAI, Sora 2 is necessary. It needs a standalone product to truly productize the technology and land it with appropriate boundaries. This isn't just a technical problem, but organizational mechanism and product system construction. It needs an entire operational flow from technology to product — SOPs, organizational structure, team collaboration — these are what truly matter for OpenAI.
👦🏻 Koji
When I saw that article headline saying "The world no longer needs an AI Douyin," my first reaction was: actually OpenAI never intended to build an AI Douyin at all, right? The proposition itself is externally imposed on them.
👦🏻 Minghao Zhuang
That day Sam mentioned that OpenAI is becoming increasingly complex. Internally it's now like three or four companies operating in parallel: a product company, a technology lab, a technology infrastructure company, and possibly incubating hardware-related businesses. It's no longer the simple startup it once was — now valued at $500 billion, with at least four clear entities plus potential new businesses. Recently Sam has been busy with that infrastructure company. Its current capacity is sufficient to meet OpenAI's own needs, but if it truly invests over a trillion dollars, this company's capabilities could even overflow to support other enterprises.
👩🏻 Ronghui
Did you see after that OpenAI launch event, many well-known companies on X were showing off the commemorative badge they received, like a medal? Companies like Figma were all posting about it.
👦🏻 Koji
Right, the "how many tokens used" badge. They're just so good at marketing. I remember last year we actually planned an episode around when Sam Altman was getting flamed as a "marketing hack" — we wanted to discuss how to learn marketing from Sam Altman.
👩🏻 Ronghui
I think the Sora 2 videos themselves were an extremely well-calibrated communications experiment. Neither overstepping nor flashy, but making full use of the company's existing distribution resources. Especially using the founder as an IP — very Elon Musk in that way. Mock me all you want, I don't care.
👦🏻 Minghao Zhuang
Right, he's thought it through completely.
👩🏻 Ronghui
And this approach strengthens brand lock-in. Every interaction users have with Sam or OpenAI deepens that emotional connection.
👦🏻 Koji
Yes, we've been talking about OpenAI for quite a while now. Almost without realizing it, the world is still being shaped by them. No matter how much skepticism, controversy, or brief pessimism the outside world experiences, you still have to admire their sense of rhythm and creativity.
But I also want to talk about Google. This year, Alice and their podcast did a parallel episode on "Google DeepMind's comeback." They interviewed DeepMind executives who said that whether it's Nano Banana or Veo 3, Google's return signals are clear. Especially on the world models front — they released Genie, and when I saw it I genuinely felt this electric jolt through my body. Absolutely incredible. It's just that Genie is still quite far from commercialization, so there hasn't been much follow-up discussion. But I'd love to hear your take on this world model path, Minghao.
👦🏻 Minghao Zhuang
If we push this question back a bit, it actually approaches philosophical territory. Ever since natural language models emerged, many technologists and even philosophers have believed that superintelligence arrives when machines master language — this uniquely human capability. That's one school of thought.
The underlying logic of language models holds that language is the core structure of human civilization; all wisdom and creation build upon it. But then a branch emerged — Coding. Some see programming languages as merely a subset of language, while others view them as an independent gateway to a new world.
Further along came multimodality. People realized that language, however powerful, cannot fully express the sensory world — especially vision. So multimodal models progressed from speech, images, video, 3D, to today's visual models, gradually converging into what we call "world models."
DeepMind executives mentioned in interviews that they see world models as another main path to AGI. Using a Texas Hold'em analogy: language models are one main table, world models are another. DeepMind describes world models as "the womb of the world" — a self-consistent generative system. If this system can generate worlds governed by real physical rules, and do so faster than human imagination, it truly gestates a new form of intelligence.
Of course, this is an extremely idealized technical vision, but DeepMind has been pushing in this direction all along. In fact, they were working on this before the language model explosion.
👦🏻 Koji
Right, they've always been doing leading-edge innovation.
👦🏻 Minghao Zhuang
Exactly, and Fei-Fei Li follows the same logic. Her startup is also building world models. So many believe that if world models can truly become another main path, they might achieve AGI in a different way. From a long-term perspective, this ideal state is genuinely breathtaking.
The previous generation of AI companies mainly solved vision problems, and vision's commercial scenarios (images, video, etc.) are already so massive that there's no need to open new territory. Therefore, when world models combine with existing scenarios like gaming and video, they can almost instantly embed into application chains. It's just that the technology hasn't reached its tipping point yet — there hasn't been a GPT-3.5-level breakthrough moment.
World models might even be more difficult than language models.
👦🏻 Koji
Besides DeepMind and Fei-Fei Li's World Lab, who's doing world models domestically?
👦🏻 Minghao Zhuang
Tencent's Hunyuan is working on it. They just released version 0.1, still very early — can only transform a photo into an interactive 3D form. The image quality and pixel clarity are still quite primitive. But the logic is consistent: competition in images and video is already white-hot, while 3D and world models still have huge room. The main players here will be gaming companies, since they have natural demand scenarios.
The Year of Agent, Then What?
👦🏻 Koji
I heard an interesting data point the other day. Bambu Lab now has its own platform where users can generate 3D models and print them. They can call Hunyuan, Mast, and several other models, and their call volume on both Hunyuan and Mast ranks at the top.
Their main customers are gaming companies on one hand, 3D printing users on the other. So multimodal entrepreneurship remains a very hot track. This year we've also seen companies like Higgsfield emerge, and the founder of Hix AI also did Pollo AI — their data and revenue are both growing fast.
I chatted with a HSG-backed company recently. They just launched a video tool product. I asked what their differentiation was, since there are already many similar products like VEED, and they replied that differentiation doesn't matter that much — the market is too big.
They gave an example: TikTok adds 1 million users globally every day, and 12% of them tap that plus button in the middle — so roughly 120,000 new video creators daily. However grassroots they may be, they all need video tools. Currently in the video tools space, there are already 20 companies with ARR over $20 million, many of which we've never heard of. They might just solve one specific problem, or focus on one country. So multimodal remains a very promising domain.
When reviewing 2025, besides DeepSeek R1, we have to mention another memory belonging to all Chinese people — the launch of Manus. "Agent" became the keyword of the year. Crossing's opening podcast episode this year was titled "The Year of Agent". That episode with Yusen was a bit risky for us, because if this year didn't turn out to be the year of Agent, we'd have egg on our faces. But fortunately, Agent development has indeed met expectations.
Now there's more and more discussion: Should agents be general-purpose or vertical? Will ChatGPT eat all agent companies? These are all typical questions. I'd love to hear your thoughts, Minghao.
👦🏻 Minghao Zhuang
I'll start from the L1 to L5 model classification. L3 corresponds to Agent. When everyone has achieved L2 — the ability to understand and generate language — the natural next step is L3. L1 to L2 still solves language problems, but L3 begins to involve "behavior" — not just conversation, but actually getting the model to execute tasks. Whether on computers, web pages, or in databases, it needs to "act." This makes things much more complex.
So even though we're calling this the year of Agent, it could last for many years, like VR or autonomous driving — we might still be in the "year of Agent" phase five years from now.
Because while Agent is a milestone node, there are many levels and sub-stages within it, with different levels corresponding to different experiences and needs. We're therefore seeing new bifurcation — general-purpose agents and vertical agents each have their own paths. Certain vertical scenarios may more easily achieve idealized agent deployment.
The emergence of Agent gives non-model companies an entirely new paradigm for landing. Previously, people wanting to build in AI faced model vendors' monopolies with nowhere to start, only able to "wrap" — not a pejorative, but because what they could do was so limited. But today, the Agent ecosystem is expanding rapidly, complex and diverse. If we analogize to the early internet, the Agent space hasn't even standardized "protocols" yet. Major companies are all competing for protocol-setting power — whether Claude, Google, or OpenAI.
Once some de facto standard forms, the protocol becomes established. Then, when moving from language to behavior, we need to build massive amounts of "scaffolding" (Infra) to be compatible with existing internet systems. The question is: keep using browsers, or go pure API? Build on old systems, or rebuild from scratch? These choices determine differences between companies.
So various types of companies emerge: protocol layer, Infra implementation tools, underlying infrastructure, even those focused on memory systems. The entire ecosystem becomes extremely complex because of this shift.
👦🏻 Koji
Speaking of memory infra, that's indeed a critical point. When we interviewed MeMU before, we noticed that Infra around "memory" had already become dizzyingly abundant. I later chatted again with MemoBase's founder, who helped me map out the entire technical landscape. Only then did I realize how completely different routes companies are taking. He said there's simply no unified technical standard right now — everyone is using their own approach.
👦🏻 Minghao Zhuang
We're also exploring related directions ourselves. Because our main business is social, and in AI social, memory systems are extremely important — whether individual memory or interactive memory between people.
Our technical lead has been quite troubled by this: build in-house, use open-source solutions, or choose established API providers? Even this choice is hard to settle.
👦🏻 Koji
The answer is blowin' in the wind.
👦🏻 Minghao Zhuang
Yes, there's genuinely no standard right now.
👦🏻 Koji
Chunxiang Zhao came to share at our Beijing AI Open Mic recently, also spoke for 10 minutes at Crossing. His topic was a new project — an IM product. But the focus of that talk wasn't the IM itself. It was the memory system he built in-house for this IM. Because he felt existing memory systems on the market weren't good enough — either unstable or poorly suited to his scenario — so he decided to build his own.
👦🏻 Minghao Zhuang
Yes, and that's just "memory" alone. I remember Ant Group open-sourced a massive ecosystem map not long ago, covering Agent frameworks, Agent Memory, Agent Infra, and other segments. Over the past year, the fastest-growing directions in the entire open-source ecosystem have almost all been related to these modules.
Because everyone has genuinely reached this point — to build an Agent, you have to solve these underlying problems. Many entrepreneurs now face the same predicament: the ecosystem is too new, with no mature standards or tools, so they end up building everything themselves. It's like an iceberg — what we see is just that small tip labeled "Agent," while the massive system underneath is where the real difficulty lies.
👦🏻 Koji
And only during a chaotic period like this can startups establish their own unique competitive advantages.

👩🏻 Ronghui
Speaking of which, I have a curious question. Do you ever get the feeling that while OpenAI talks about Agents too, they've been remarkably restrained in their later approach?
👦🏻 Minghao Zhuang
I think it's because in their decision-making hierarchy, Agent isn't the highest-priority direction. They still lean closer to the traditional internet product manager's mindset, focusing on "demand fulfillment" and "product form." And the Agent that people talk about now has been somewhat narrowed in scope — its interaction patterns, output formats, button styles, UI design — almost boxed in by the imagination of the past year.
OpenAI is making adjustments too, like launching their own Agent, reshaping Operator into its current form, and integrating Deep Search. Actually, they didn't fully figure out what an Agent should look like at first either. But after Manus launched in March, that prototype gradually became clearer.
I think Manus's greatest significance is that it was the first time users actually "saw" what an Agent product should be. Whether in interaction logic, interface design, or overall experience, it provided a concrete reference point.
👦🏻 Koji
At this stage, the dividend of occupying the "mindshare synonym" is enormous. So it's not surprising that Manus recently announced $90 million in ARR. Of course, plenty of people later joked that they could replicate a Manus in three days. But you can't deny that a truly successful product relies on far more than technology alone — it's a manifestation of combined capabilities: timing, attention to detail, and brand building.
👦🏻 Minghao Zhuang
Right, and this brings us back to that question — the brand war for Chatbots is basically over, but the brand war for Agents is still ongoing. Manus just seized a very strong first-mover advantage, and seized it very solidly.
👩🏻 Ronghui
Actually, during the summer, a lot of people were saying Agents were cooling off. Did you feel that?
👦🏻 Minghao Zhuang
If anyone remembers, GPT-5 launched in August, and so did that Claude version. August was the month when several leading US companies concentrated their releases, while July became almost the most intensive open-source month for Chinese companies — new models launched every single day. It was a genuine "Crazy July."
Minimax even staged a "launch week," releasing something new for seven consecutive days. Domestic players like Zhipu AI and Moonshot AI were all pushing hard during that period too. So attention in July and August did swing back to "pure model competition." Many companies started pushing products, acquiring users, and doing implementations; Infra companies also took the opportunity to raise funding and expand, even pivoting in other directions.
👦🏻 Koji
True, the attention wasn't at the viral, screen-flooding level of Manus's launch. But to say it "cooled off"? Not at all. Because shortly after, Manus quickly announced $90 million ARR and introduced the concept of "RRR" (Realized Recurring Revenue) — really just popularizing a more rigorous revenue metric. And $90 million in scale already exceeds most high-growth companies we know.
Just last week, a16z Speedrun held its 005 Demo Day, with 58 startups each presenting for two minutes in LA. Many attendees were saying on Twitter and Jike that the quality of this batch even surpassed YC's Demo Day. I think that may well be true — a16z has more money, and can attract more mature founders.
A friend of mine recorded audio from the event, and I had Gemini listen and summarize for me. These companies roughly fell into three categories: the first was Agent as a Service — solving specific problems across various verticals. So Agents absolutely haven't cooled off; they've just entered a deeper, more fragmented phase. Instead of relying on viral presence, they've already permeated into all kinds of entrepreneurial scenarios.
👦🏻 Minghao Zhuang
Around the same time, a16z also released a semi-annual list — a startup revenue ranking produced in partnership with a payment data vendor. The list tracked which AI companies these startups were spending money on. Four Chinese companies made the cut: CapCut, Keling AI, Manus, and Genspark.
👦🏻 Koji
Looking back, since Siri launched in 2012, phones have basically all had voice assistants. But compared to today's Agents, there's still a generational gap.
👦🏻 Minghao Zhuang
Yes, this is actually a continuous evolution along the same lineage. First, voice model capabilities have significantly improved. Second, the progress in voice models this round is "end-to-end." If anyone remembers, one of the most stunning demos of a certain GPT version was the voice conversation, the phone call scene — that demonstrated the power of end-to-end voice models.
If we view the entire internet ecosystem as a complete system, existing applications provide mature scenarios and experiences, while phone manufacturers possess user identification and scenario-grounding capabilities. Therefore, they won't and can't complete all functions themselves; instead, they choose to partner with service providers who already have mature scenarios.
Just like how OpenAI's ChatGPT partners with external vendors — the logic is exactly the same. In the Chinese context, the advantages of this ecosystem collaboration are even more pronounced. Because China's mobile internet ecosystem is extremely mature, with extremely broad scenario coverage. As long as front-end demand analysis is done well, there will definitely be quality service providers coming forward with solutions.
And the implementation of these services might be through Agents, or through search, or traditional lists. But the core point is: the formation of this entire ecosystem is built precisely upon the rich existing foundation of China's mobile internet.
Open Source: An Ecosystem with Chinese Characteristics
👦🏻 Koji
This year there's been a noticeable shift — over the past two years, AI has been accelerating its landing on device endpoints. Combined with the rapid development of Agents, device manufacturers generally believe that integrating Agent capabilities is an important direction for the future, and they've been placing bets accordingly.
We've already discussed large models, multimodality, and Agents — that constitutes a fairly rich system. Next let's talk about ecosystem, starting with the "open-source ecosystem."
Open source has become a consensus in China from top to bottom, and a national-level strategy. Recently Shanghai even introduced policies encouraging companies not only to create open-source products but to build open-source communities. The policy even offers cash rewards up to 5 million RMB for projects with sufficient influence and traffic in the open-source domain. The prerequisite for all these changes is clearly the success of DeepSeek R1. It showed everyone that for Chinese companies in global competition, open source is a viable and competitive path.
So Minghao, what's your take on how open source has developed this year and what significance it brings?
👦🏻 Minghao Zhuang
Actually, as early as last year, a reporter asked Sam Altman about his views on "open source" versus "closed source." His answer at the time was that the most leading models would definitely be closed source, because this requires enormous capital and resource investment. But with the rise of DeepSeek R1, that view began to be shaken. Now the industry generally believes that open-source model capabilities are rapidly catching up, and that future top-tier models will inevitably include open-source ones.
Even in the "State of AI Report" published just a few days ago, there's a prediction: by 2026, there will definitely be an open-source model that reaches number one globally for a period of time. Additionally, after GPT-5's release, some of its models were open-sourced too — OpenAI has also been forced to participate in the open-source system. When open-source models' foundational capabilities are strong enough, the performance gap between open and closed source is no longer significant. And given the massive cost differences, open source has greater advantages in developer ecosystem and application-side scalability.
Now evaluating whether a model matters isn't just about Benchmark scores; what's more important is: how many people use it, its compute consumption, how many vendors it covers, and its penetration in the ecosystem. Because even the most powerful model means nothing if nobody uses it. From this perspective, open source provides model vendors with an alternative path to gain recognition, accumulate attention, and continuously improve their models through community collaboration.
In the past, people thought open source "doesn't make money" — it was idealistic technical romanticism. But now open source and commercialization aren't contradictory. Take DeepSeek: it's open source, but its API is also paid. B2B enterprises accessing its API need to pay. Although DeepSeek's API costs may be dozens or even hundreds of times cheaper than OpenAI's, as long as its cost structure is controllable, it can establish an independent business model this way. More importantly, the existence of the open-source community allows more developers and vendors to jointly improve the model. Theoretically, the more it's used, the faster the model improves.
From a more macro perspective, global AI competition today is essentially competition between China and the United States. Beyond the direct US-China rivalry, the bigger question is: how do other countries choose? They all need models — so do they choose OpenAI, or Chinese companies, or build their own? Japan, African nations, Southeast Asia, the Middle East — they all face this decision. Therefore, open source has also become, in a sense, a "soft weapon" with strategic significance.
👦🏻 Koji
Exactly. Open-source models tend to more easily gain goodwill globally. They allow more people to deploy models locally at lower cost, thereby acquiring users faster. At the same time, it naturally solves the "trust" problem — many worry about data leakage or misuse, and open-source models make everything transparent and verifiable.
👦🏻 Minghao Zhuang
For instance, during the recent "Double Eleven" shopping festival, I noticed one laptop model sold out almost instantly. The reason? It was especially well-suited for running large models locally — plenty of RAM and VRAM, yet relatively affordable. For people working in large model research, this kind of hardware is a "must-have."
👩🏻 Ronghui
I completely agree with your point about "marketing." Take Liang Wenfeng of DeepSeek — he never gives media interviews, yet he has his own way of communicating with the tech world.
👦🏻 Minghao Zhuang
It's his identity as "first author."
👩🏻 Ronghui
Their papers inspire optimism and confidence in where technology is headed. In a world where various forces beyond our control limit communication, there still exists a world that can connect through its own distinctive language and methods.
I remember after DeepSeek released R1, I watched Lex Fridman's five-hour podcast. He analyzed why R1 performed so well with remarkable objectivity and calm — barely any political bias. It made you feel that people in the tech world still believe that if a product is good enough, it can be evaluated and recognized on its own merits.
👦🏻 Koji
What notable breakthroughs have you seen at the application layer this year? Beyond the Agents we discussed earlier, what other keywords or applications would you consider defining moments in this year's tech memory?
👦🏻 Minghao Zhuang
Coding is definitely unavoidable, especially AI coding. This year saw significant progress across front-end, back-end, database, and other specialized domains.
Another area is vertical-domain Agents. General-purpose Agents have become the dominant approach — both large model companies and startups are building them. But in vertical sectors like legal services, several companies raised substantial funding this year. These companies require extensive vertical industry knowledge and workflows, plus they must meet privacy and data security requirements. Not just one or two — maybe five, six, or even seven or eight companies in this space have raised money and are developing well.
Finance is similar, segmented into primary markets, secondary markets, insurance, banking, and so on. Marketing has always been the most competitive domain, covering online and offline, search engines, video, images, text, email, and more. While these companies aren't large in scale, the sector is thriving.
I think there's another area that doesn't get mainstream attention but has been steadily developing — social and companionship. a16z's latest rankings show 10 companies in the top 50 on web, and 12 or 13 on the app side, yet almost no one discusses them. Perhaps because social has already gone through multiple generations: 1.0, then 2.0 with added interactivity and immersion, then 3.0 with scenario-building, and it's still evolving further.
Starting from Q3 this year, several domestic companies doing AI dynamic comics (AI-animated or motion comics) raised significant funding in the primary market. The ceiling for the anime sector is quite obvious — the previous mobile internet generation already tried it with virtually no results. AI adds more possibilities, but even pushed to its limits, there won't be massive breakthroughs. It's simply that advances in image generation, inference models, improved usability, combined with content demand from anime platforms, overseas expansion narratives, and short drama synergies, have "squeezed" this sector into receiving more attention than expected.
This also reflects capital market dynamics — when first-tier sectors are overcrowded, entrepreneurs, VCs, and investors all need new "second-tier tables." The first-tier tables simply can't accommodate so many people anymore.
👦🏻 Koji
So you're not bearish on AI dynamic comics?
👦🏻 Minghao Zhuang
Not at all — we actually invested in one. It's just that its window of market attention won't last very long, and the recognition window is limited, so you have to move fast.
👩🏻 Ronghui
Do you think anything similar to Cursor has emerged domestically?
👦🏻 Minghao Zhuang
Cursor or Cursor-like applications — perhaps two dimensions. One is coding, which everyone has; but another is truly defining a paradigm or category like Cursor did, which I think is quite difficult.
👩🏻 Ronghui
Why?
👦🏻 Minghao Zhuang
Mainly because China, during the mobile internet era, developed more mature pure-toC product design and operational capabilities than the US, so people expect an explosion of pure-toC applications. But so far, we haven't seen this hypothesis validated.
Looking at the rankings, it's still the apps we mentioned before. Even when "new" names appear, look closely and 90% are traditional apps with just 10% AI functionality added, yet they get placed high on "AI application rankings." This is hard to agree with — many of these companies have been around for years, they're not really new.
👦🏻 Koji
Like which ones? Haha, I know who you're thinking of.
👦🏻 Minghao Zhuang
Can't say that! Would offend people, haha.
👩🏻 Ronghui
Those American apps always have headlines like "reached X ARR in Y time."
👦🏻 Minghao Zhuang
Right, that might be the American trend. But in China we can't copy that, because we can't talk about ARR. So what can we evaluate? Basically just user numbers. But in today's China mobile internet market, even counting web, achieving high user numbers is extremely difficult. This isn't an AI problem — it's an industry-wide problem.
Quick Mobile's first-half statistics categorized all AI applications into three types: standalone apps, web, and plugins. Among standalone apps and web applications, three-quarters had negative user growth, only one-quarter growing. Meanwhile about two-thirds of plugins were growing.
But the question is, who wants to just be a plugin? That's the reality. Why are plugins growing? Because under China's current mobile internet system, all entry points are already occupied — like keywords. If you don't do plugins, trying to carve out a new battlefield on your own is incredibly hard.
Returning to those companies we mentioned earlier, like AI image communities — they've all raised funding, some two or three rounds, even three or four. The founders are excellent, doing the right things. But even so, explosive user growth in a short time is hard to achieve. That kind of super-breakout trend is virtually impossible today. During the mobile internet explosion era, new apps could hit a million DAU in a week — that basically won't happen now. Even if it does, it's like a shooting star, gone in a flash.
Secondary Market Frenzy, What About Primary?
👦🏻 Koji
Right now the US has many exaggerated funding stories — almost every week some company you've never heard of raises $50 million. Many of these are toB. But in China, this rarely happens. Partly because SMEs' willingness, ability, and awareness to pay are insufficient; partly because our mobile internet is too mature.
Mobile internet today is a place full of "castles" — WeChat is a castle, Douyin is a castle, Xiaohongshu is a castle. In the US, the web is more open, encouraging various tools to collaborate and achieve 1+1>2 effects. This difference is clearly reflected in entrepreneurial ecosystems and opportunity choices. I think we're more passive now.
👦🏻 Minghao Zhuang
Indeed, the familiar patterns we used to know are now too hard to replicate. We need new trends, but can't completely copy American models because the underlying foundations differ. So what should our trends look like? This isn't just entrepreneurs' problem — it's investors' too: what should we invest in?
For dollar-denominated VC funds right now, AI dollar investing is genuinely very difficult. You invest in a large model project, follow a few rounds, and then it's over — just wait for the outcome. Everyone says look at Chinese companies for AI applications, but if it's an overseas expansion team, you face all sorts of complex issues. If you invest in a purely domestic team, you have to ask: what are you expecting? What's the expected exit? These questions are virtually unsolvable.
Of course, I don't do pure VC now, so I can say this lightly. But putting yourself in their shoes, these problems really make investment work hard to carry out. So you have to set aside these unsolvable questions, focus on the founder's background and direction, communicate seriously, judge whether they're worth believing in, then decide to invest. Don't overthink, because thinking won't clarify anything.
👦🏻 Koji
A few days ago I met a managing partner at a fund and asked how he views the current dollar VC environment. Everyone generally feels it's tough, but he said he asked himself one question: do you still believe in VC's underlying narrative — that tech companies need patient capital in early stages to help them build from 0 to 1. If you still believe, keep doing it.
👦🏻 Minghao Zhuang
This reminds me of yesterday's Nobel Prize in Economics announcement — one of the laureates' research showed that what truly drives rapid economic growth is "creative destruction." This kind of innovation is tightly bound to VC's development logic.
If you believe in this underlying logic, of course VC still needs to exist. But the problem is, this is an overly idealized, ultimate answer. The logic isn't wrong, but when implemented in reality, especially domestically, there are too many constraints and limitations. Investors can only search for solutions within limited space.
This is also the scale that many investors who still believe in technology, believe in products, believe innovation will change the world — keep weighing back and forth in their minds.
👦🏻 Koji
What about beyond AI? For instance, embodied intelligence — there's been heavy investment this year too.
👩🏻 Ronghui
AI hardware is actually very hot.
👦🏻 Minghao Zhuang
Yes, AI hardware is indeed hot. We can see the shifting backgrounds of funded founders. Initially it was product managers from ByteDance and other major internet companies who got funding; earlier still it was research backgrounds — lab professors or researchers.
But later people found both types had limited results. So investment shifted toward industry founders — entrepreneurs from companies like DJI, Dreame, Roborock, Insta360.
👦🏻 Koji
I saw one media headline: financial advisors have already opened offices next to DJI.
👦🏻 Minghao Zhuang
Exactly, the logic is clear: investors have already funded everyone in their familiar circles. Why invest in hardware? Because compared to pure software, hardware is more tangible, more visible — there's physical product, there's a revenue model, giving investors more sense of security. Plus China's supply chain advantages make it easier to produce concrete results.
But even so, many projects still underperform. So another round of logic emerged — go find founder backgrounds that make investors "more trusting, able to sleep at night." This logic all sounds correct, but from another angle, we're actually "talking without skin in the game." Entrepreneurs or investors on the battlefield can't survive by summarizing patterns — they must make decisions amid uncertainty.
👩🏻 Ronghui
I listened to Uncapped's podcast yesterday — Jack Altman interviewing Thrive Capital partner Vince Hankes. He told a story: they spent 18 months researching one company before investing. That would be impossible in China.
👦🏻 Minghao Zhuang
18 months? In China that's already four cycles.
👦🏻 Koji
Vince came to China two months ago. He met several of the companies we just mentioned, very interested in China, arranged an extremely packed itinerary.
So here I am in 2025, and I'm genuinely optimistic about China.
👦🏻 Minghao Zhuang
That optimism actually carries a bit of philosophical weight — the extreme turns back on itself. When any domain pushes to an extreme, a new equilibrium eventually emerges. Take gaming investment. After 2015, Chinese gaming largely fell off the VC map entirely. Yet over the past two years, overseas VCs have started paying attention to early-stage Chinese gaming teams again. Keep in mind, Chinese gaming teams haven't been in VC crosshairs for nearly a decade.
👦🏻 Koji
What's their logic? How do they exit?
👦🏻 Minghao Zhuang
They're not starting with exit — they're starting with cost. Compared to overseas teams, Chinese gaming teams are dramatically cheaper — and I mean comprehensively cheaper. Meanwhile, gaming M&A overseas has stayed extremely active. So they're combining the two: develop content cheaply in China, then monetize through overseas acquisitions. This is something Chinese VCs simply can't pull off.
👦🏻 Koji
Actually, I think Vince coming to China, and overseas VCs looking at China more broadly — the world is paying very close attention to China.
We've talked a lot about primary markets, but recently it's probably the secondary market that's grabbing more headlines. When the OpenAI and AMD news dropped, AMD — a company of that scale — jumped 40%. The day before we recorded this podcast, the OpenAI and Broadcom collaboration was announced, and Broadcom's stock immediately rose 10%.
So there's a lot happening in public markets, and the mood is genuinely optimistic. Minghao, you've been in frequent conversation with sell-side analysts and fund managers over the past year — how would you look back on this year?
👦🏻 Minghao Zhuang
I want to frame this through the reports I've produced. So far this year I've published six reports, with a seventh in preparation.
These include last year's year-end review, DeepSeek, Manus, Agent analyses, and my Q2 and Q3 market summaries this year. It wasn't until my fourth report that I even raised the question "Is this a bubble?" By the fifth version, I had two pages on bubbles. By the September Q3 report, six pages were devoted to "Are we in a bubble?" That shifting proportion shows the conversation is heating up.
In Q2, some voices started sensing something off; by Q3, I was citing Silicon Valley Bank's report most heavily — they did extensive qualitative and quantitative analysis, including curve comparisons to the 2000 dot-com bubble. SVB was already researching this back then, and now more and more people are adopting similar analytical frameworks.
I still believe we've genuinely, unconsciously arrived at some kind of limit. Take the "Magnificent Seven" — several now sit at $3-4 trillion market caps, generating hundreds of billions in annual operating cash flow. OpenAI's valuation: $500 billion. These numbers mean they can do anything.
The market's no longer satisfied with retail investors pushing prices higher — that game isn't exciting enough. Meanwhile, OpenAI has reached an inflection point. I think Sam had an epiphany at some moment: OpenAI's capital-layer operations were mainly about fundraising, selling secondary shares, pumping valuation, then using that capital or equity to invest in startups — like the recent acquisition of the company founded by Apple's former designer. Valuations doubling every 6-9 months, each round raising $20-30 billion — historically unprecedented, but still within a controllable, linearly extrapolatable range. At least someone was buying, contracts were signed.
But Sam seems to have glimpsed another path. He's lately being compared to a "cloud-to-everything" play — Google owning cloud, models, products, and infrastructure. If AI eventually permeates everything, and Google's worth $3 trillion, what should OpenAI be worth? How far out can growth expectations stretch? It used to be six months or a year; now it's five years. So recent contracts are almost all 5- or 10-year terms.
The irony: technology advances year by year. We've seen countless analyst, bank, and brokerage reports — nobody writes 5-year technology forecasts. But now Sam says he's compressing the next five years of expectations into this one bet. OpenAI at $500 billion — that magnitude is globally rare. Even pre-IPO, it can move entire markets. So the flood of news we've seen this past month is precisely the result of this valuation dynamic.
Back in 2024, Sam talked about building this to $7 trillion. That sounded insane then — OpenAI was only valued at $180 billion. $7 trillion exceeded almost anyone's comprehension. But now, when we aggregate the five-year forward expectations for OpenAI, NVIDIA, Microsoft, Google, and Oracle, trillion-dollar magnitudes are being openly discussed.
Whether it reaches $7 trillion, nobody knows. But at least "trillion-scale" is now in conversation. All narratives, all expectations are compressed into the next five years. Whoever takes investment must max out their own five-year projections too. And the competitive landscape is so intense that no one can afford to sit out. AI is no longer just technology or industry — it has become the market itself.
AI today isn't merely a sector; it's the entire system. "Too big to fail" doesn't even capture it anymore. The ecosystem is so complex, everyone's swept up in it, the scale so vast that rational assessment has lost meaning. I keep emphasizing: We may have arrived at a limit that no one has fully recognized.
As of today, the largest IPO fundraising by any listed company globally was roughly $25 billion.
👩🏻 Ronghui
For a single company?
👦🏻 Minghao Zhuang
That was Saudi Aramco — a world record, yet already smaller than OpenAI's single funding rounds, even smaller than many cloud computing companies' individual raises. If OpenAI went public today, how massive would it be? How much capital would it need to raise? The entire stock market would likely be torn open.
I saw an interesting data point the other day: In the Web era, the most cash-burning public company was Amazon, which burned $2 billion pre-IPO. In the mobile internet era, the biggest burner was Uber, at roughly $40 billion — 20x Amazon. Does that mean OpenAI needs to burn 20x Uber, or $800 billion, to go public?
Now that figure no longer seems fantastical — it may well be reality. It's been positioned there. For whatever reasons, this era has arrived here, carrying America's enormous technological momentum, Silicon Valley trends, and financial trends, rolling to this stage — it can only keep moving forward, no alternative.
Continuing on, countless people are drawing comparisons: to the dot-com bubble, the railway mania. Of course, nobody mentions tulips or Bitcoin anymore — that's not the same category. We're talking about genuine industrial-revolution-scale historical transformations.
The hottest analogy recently is the internet. When the dot-com bubble burst, Nasdaq-listed internet companies nearly collapsed entirely, but the underlying fiber infrastructure remained, laying groundwork for two decades of internet development. Yet the reality is — the companies that laid that fiber mostly died. Does that mean today's data center builders might die too?
The future could certainly be better, but there's a critical difference: fiber, once laid, still works fifteen or twenty years later. But data center GPUs being deployed today depreciate over three years, and may be obsolete two years after that. The depreciation cycle creates enormous pressure.
The bigger problem: the entire industry's massive capital deployment rests on the extraordinary revenue power of a few giants. Companies like Meta, Google, NVIDIA, Microsoft — earning hundreds of billions annually, so investing hundreds of billions is manageable. But beyond these few, almost everyone else in the ecosystem is straining. Cloud service providers, even Elon Musk's xAI, now need new funding sources and have begun using debt financing.
When analyzing bubbles, equity bubbles are one thing; debt bubbles are far more dangerous. The subprime crisis happened because debt is rigid — miss a payment and everything collapses. So when we strip out those few giants, the underlying risk profile is actually extremely high. Companies like CoreWeave already have pretty ugly balance sheets. They started as mining operations, running on highly leveraged classical financial logic. While not uncommon in the Web3 world, this is spreading into a broader trend.
There's another fascinating signal. As of mid-October 2025, the two best-performing stocks in the S&P 500 were neither GE nor any energy utility, nor gold miners — they were two hard drive manufacturers: Seagate and Western Digital.
Why? Because people realized the data center construction frenzy has arrived. First wave benefited NVIDIA, then cloud computing companies, then transformers, cooling, power — each sector got its turn. Finally everyone noticed: you also need hard drives, you need storage. But hard drives are constrained by NAND flash supply. This is why Sam Altman recently flew to Korea and Japan, visiting SK and Samsung — to secure storage supply.
Returning to that earlier point: we've genuinely, unconsciously arrived at limits across many domains. Power, cooling — discussed endlessly already. Now even storage has become a constraint. And this is another nearly monopolized industry.
So you see, when all these factors stack together, it becomes impossible to distinguish who's responsible, which company's fault it is. It resembles a massive conspiracy — not an active one, but something that's been pushed and rolled to this point. By the time we realize it, there's no turning back, only forward. That's the whole story.
👩🏻 Ronghui
Earlier we mentioned quite a bit about hardware and compute progress in Europe and the US in 2025. But at the same time, domestic China has also made massive investments and achieved significant results in related areas over these past few years. Anyone in the industry surely remembers how explosive this summer's World Artificial Intelligence Conference and World Robot Conference were — tickets were nearly impossible to get.
Let me add one detail: when researching, I asked AI to summarize compute-related developments for me. Several reports all highlighted the same point — China's "East Data, West Computing" initiative built out an intelligent computing network exceeding 30 billion FLOPS in 2025. This national project launched in 2022; the core concept is shifting eastern China's data processing demands to the west, optimizing nationwide compute resource allocation.
What does 30 billion FLOPS mean? The AI gave me a vivid analogy: if your smartphone can perform 100 calculations per second, then 30 billion FLOPS equates to 300 trillion smartphones working simultaneously for one second. In other words, this level of compute is sufficient to support ultra-large-scale AI models, process massive urban datasets, and simultaneously deliver internet services to billions of users.
Another major hot topic is naturally robotics, or what people commonly call "embodied intelligence." China has been the world's largest industrial robot market for the past 12 years. The World Robotics Report published in September noted that China still ranks first globally in both new installations and operational stock of industrial robots (based on 2024 data), with 2025 expected to see further growth.
Another key figure from the report: by 2024, domestic Chinese manufacturers' sales surpassed foreign suppliers for the first time, with market share climbing from 28% a decade ago to 57% — nearly doubling. This means China is not only the largest market but has also achieved a qualitative breakthrough on the manufacturing side.
I remember we previously had a podcast interview with Zhelun Zhao, co-founder of Vbot. He happened to be in Beijing attending this year's World Robot Conference. He mentioned he was heading to the "Robot Games" that evening and shared some vivid observations. One thing stuck with me — he said this "Robot Games" is like F1 in car culture. Its significance lies not just in competition, but in cultivating cultural soil for the entire industry. The establishment of this culture could have profound effects on the future industrial ecosystem.
He also mentioned a detail: at the conference, many children were interacting with robots. He remarked at the time — our generation grew up with computers and phones, while today's children may become the generation that grows up alongside robots.
From a longer-term perspective, this may well mark the beginning of a new era. It also echoes the question we previously asked Minghao — "What are the things that, looking back ten years from now, we'll still find memorable about this year?"
Perhaps it's precisely these changes that make one truly feel we are witnessing a historic starting point.
👦🏻 Koji
We started with large models, multimodal, Agent, then moved to secondary markets. We began talking about entrepreneurship, opportunities, technology products, and it ultimately became a flowing conversation with real historical depth.
👦🏻 Minghao Zhuang
I'm Sima Qian, then.
👦🏻 Koji
Hahaha, this was a fascinating episode. Thanks Minghao, thank you! Looking forward to our recap same time next year.
👦🏻 Minghao Zhuang
Thank you!
