Large Language Models Are Still in Elementary School — Don't Rush Them Into the Workforce | A Conversation with ZhenFund's Yusen Dai: What Stage Is AI At Right Now?

This is a crossover episode between the ZhenFund podcast *"Seriously Speaking"[1]* and the Crossing podcast[2].

This is a crossover episode between ZhenFund's "True Words[1]" podcast and "Crossing[2]" podcast.

A year and a half ago, ZhenFund and Yusen were among the earliest fund and investor to start paying attention to AI. He acted decisively, responded quickly, and rapidly accumulated a wealth of frontline insights, investing in multiple leading AI companies including the large model company Moonshot AI.

After a year and a half of frenzy, AI has recently shown signs of cooling down, and the view that "it's shameful not to monetize" has become increasingly mainstream in China.

In this episode, Yusen and I use the internet and mobile internet as a historical reference book to analyze which page AI is on today.

At the same time, Yusen made an analogy: If today's large models are like a gifted elementary school student, as a parent, would you choose to send them out to work immediately, or continue investing in, cultivating, and encouraging them to finish their PhD before entering the workforce?

We also explore why we should give large models more tolerance and patience at this moment, and how to learn to maintain patience and optimism.

First AI Art Festival

The Fair and ZhenFund, along with other friends, jointly launched the First AI Art Festival.

This festival boasts an impressive lineup, including not only tech companies such as Moonshot AI, MiniMax, Liblib, RightBrain AI, and Tiamat, but also two VC funds — ZhenFund and Linear Capital — as well as cultural landmarks like Anaya and UCCA Center for Contemporary Art. The Fair has also invited numerous advisors and judges from the cultural sphere, such as poet Bei Dao, photographer Xiao Quan, and science fiction writer Stanley Chan.

In this episode, we introduce this festival, as well as my experience as the initiator and Yusen's experience as a judge and the person who helped make this project happen.

Recently, the First AI Art Festival will soon land an "Love, Hate, Passion, Grudge" AI image exhibition at Anaya. We will announce the specific time and location as soon as they are confirmed.

Conversation

Outline

Patience and Waiting

  • 01:46 Why do we need more patience with AI business models?
  • 03:19 How did the two super money-printing machines of the internet, Google and Meta, find their business models?
  • 08:33 "We tell the elementary school student (AI large model), why haven't you made money yet? Look at that 40-year-old uncle making so much money, and you, hahaha, no future."

Recent Reflections

  • 09:50 Reflection 1: The rapid decline in model costs
  • 12:27 Reflection 2: The rapid evolution of multimodal capabilities

Brand New Opportunities

  • 14:02 After GPT-4o's release, if humans and AI start dating, will the world become a better place?
  • 17:10 An investor's perspective on emotional companionship AI projects
  • 24:32 When will smart glasses truly take off?
  • 25:59 Achieving emotional equality for humanity — it's up to AI!

Year-in-Review

  • 28:00 What was the world and yourself like a year ago? How is it different from today?
  • 29:38 What are the most significant changes in AI investment strategy over the past year?
  • 34:08 Should large model companies also build products? Should product companies also develop their own models?

Advice for Entrepreneurs

  • 36:53 How to choose an entry point from 0 to 1?
  • 39:43 How to compete around big tech?
  • 01:07:16 How are AI entrepreneurs different from entrepreneurs of the past?

Silicon Valley and China

  • 42:47 Silicon Valley and China — how can they learn from each other?
  • 47:13 Similarities and differences in VC investment directions between China and the US?

The Future of VC

  • 51:21 In the AI era, is VC more important or less important?
  • 55:05 The story of investing in Zhilin Yang and Moonshot AI

Big Industry Questions

  • 58:04 Where is the ceiling for large models and existing technologies?
  • 59:07 In what scenarios will AI Native applications emerge?
  • 1:02:11 The most impressive Mega 7 CEO: Microsoft Satya Nadella, Tesla Elon Musk
  • 01:09:45 What are your predictions for GPT-5?

Small Life Questions

  • 01:11:23 What are your daily AI use cases?
  • 01:14:45 Using Tesla FSD in Silicon Valley — absolutely amazing.
  • 01:15:37 Do you chat with ChatGPT or Kimi?
  • 01:15:51 How to educate children in the AI era?

First AI Art Festival

  • 01:17:15 First AI Art Festival: "Love, Hate, Passion, Grudge" ❤️
  • 01:22:45 Doraemon is the most beautiful imagination of technology and AI's future

Patience and Waiting

Koji:

If last year's keyword for AI was large models, this year the key terms everyone is discussing are PMF, how to make money, and how to commercialize. The view that "it's shameful not to monetize" has also become increasingly mainstream in China. Yusen, do you agree with this view?

Yusen:

"One year in AI equals ten years in the human world." Actually, I just realized that Koji and I have known each other for almost 20 years now, back when we were special friends on Xiaonei. When Koji and I first met, we both loved the internet. At that time, we were still college students and had developed a strong faith in the internet. But back then, if you thought about the internet's business model, it wasn't actually that clear.

So at that time, whether from an investor's perspective or a career choice perspective, the internet was far from the obvious choice that people later considered it to be.

Koji:

Drawing an analogy with the internet and mobile internet, what stage is AI commercialization at now?

Yusen:

I posted something on Moments yesterday because I happened to recall that the two super money-printing machines in the internet world today are Google and Meta. Their search for business models was not smooth sailing. For example, everyone can see that ChatGPT has a huge number of users now. Of course, people question how you make money, right?

But if you recall, when Google launched in 1998, it was founded by two top Stanford CS PhDs, and they hired a large number of leading computer science PhDs from the start to build supercomputers to process massive amounts of information. From this perspective, it's quite similar to ChatGPT.

After they launched, they quickly became one of the most popular websites in the world. Everyone loved using it. But one reason people used it was because Google had no ads, so everyone found the product very useful, but it burned through massive amounts of bandwidth and computing power, while having no business model. So at that time, Wall Street kept questioning whether Google had a business model, but then it was going public, so what to do? It had to make money.

At that time, Google actually tried different business models. For example, selling search engine technology to others, or providing search engine services to others and taking a cut.

Even the hottest business model at the time was portal sites, like Yahoo and SINA Corporation back then, so they wondered if they should build their own? But later they found that none of these seemed to work particularly well.

In 2000, Google released its first business model, a more search-engine-native business model called AdWords. When you searched for a term, relevant ads would appear alongside it. Initially they charged by impressions, later switching to pay-per-click. In 2003, they released another system called AdSense that placed ads on third-party websites. Google went public in 2004, and by then the business model was relatively mature. So from its 1998 launch to its 2004 IPO — Google spent about six years after launch finding its business model.

Koji:

Actually, six years is a long time even by today's standards.

Yusen:

It's quite a long time. Then I happened to see that in 2002, The New York Times had a feature article saying that the hardest thing to search for on Google was its own business model.

Koji:

I saw your share. That article's title was very ironic: "You Can Google Everything, But You Can't Google Your Own Business Model"

Yusen:

Yes. Now everyone thinks of Google as a super money-printing machine, but just 20 years ago, nobody knew what Google's business model was, and everyone questioned it.

Koji:

What about Meta?

Yusen:

Facebook was actually the same. If everyone recalls, it launched in 2004, then launched News Feed in 2006, which established the information feed format that everyone scrolls through today. But Facebook initially had no ads either, because people didn't think putting ads on a social network was cool at all.

So at that time, they first tried letting some local merchants place banner ads, but this didn't scale much. Then they built a social gaming platform — there were some web-based mini-games on Facebook where you could top up, and they'd take a 30% cut. This generated some revenue. But these games were hot for only a year or two, flaming up in 2008 but basically fading by 2010, so Facebook needed a new business model. In September 2011, Facebook made a huge change: they transformed News Feed from chronological order — the absolute sequence like Moments — to a recommendation-based relevance ranking.

Koji:

Right, today Douyin actually launched with this kind of ranking from day one. But back then, when Facebook made this change, I remember it caused a massive uproar — users were extremely resistant.

Yusen:

Right, because I couldn't even see what my friends had just posted — maybe it wasn't recommended to me, right? I remember being absolutely shocked when I first used it, thinking, how could they do this?

You can imagine that only when a platform controls the right to recommend content can it better insert ads. So in January 2012, Facebook officially launched the News Feed advertising feature.

Everyone knows now that feed ads are absolutely a money-printing machine, right? ByteDance's money printer, Facebook's money printer. But the feed ad business model only appeared in 2012 — actually eight years after Facebook launched, and six years after News Feed launched.

Koji:

It's interesting — Facebook only got News Feed two years after launch. At first, you had to go to someone's profile page to see what they were up to.

Yusen:

So we were especially close friends: linking to each other on our homepages.

Koji:

It's easy to forget now that there was ever a time without News Feed. So you're saying ChatGPT has been around less than two years. If we use Facebook's timeline, the feature that actually makes money shouldn't even be released yet?

Yusen:

Because ChatGPT launched on November 30, 2022, so it's been exactly a year and a half. And that year and a half has been absolutely miraculous. Before ChatGPT launched, we all thought the Turing test was still a difficult thing. Look now — nobody talks about the Turing test anymore. Because obviously ChatGPT, or Kimi Chat, or any reasonably good language model can pass the Turing test. So the world has changed enormously.

But at the same time, we're now seeing skepticism in both industry and investment circles asking why AI doesn't have a business model yet? Burning so much money building large models — how does that money come back? I think a company, a product, ultimately needs to make money, but different stages have different priorities.

Previously, when we looked at mobile internet and internet business models, the technology was relatively mature. It made sense then to look at how to land the product, how to operate, how to grow, how to monetize. But AI is still in a very early stage where the technology is advancing by leaps and bounds.

If I had to make an analogy, it's like looking at an elementary school student and saying, why aren't you making money yet? Look at this 40-year-old guy making tons of money — does that mean you have no future, right?

Of course there might be some gifted elementary school students who start making money right after finishing school. But we might still think, could you maybe keep studying, get a graduate degree, and then earn more knowledge-intensive money?

As the Y Combinator founder said:

Make something people want.

You first have to build a product people want to use, then figure out monetization. I think many VCs, including us, can wait for that and are willing to wait.

Recent Reflections

Koji:

The past two or three weeks have been incredibly dense with major events. Starting from the GPT-4o launch, Google I/O, Microsoft Build, ByteDance's Doubao mega-launch releasing a ton of stuff. Among all these major happenings, was there anything that shocked you?

Yusen:

I think the first thing is that model costs are dropping extremely fast. If we look from when GPT-4 launched to the 4o API, costs dropped by roughly 90% or more. That means costs became one-tenth of what they were. And compared to when GPT-3.5 Turbo first launched, it's down to 1/40th or 1/50th. And from our conversations in Silicon Valley, people generally believe costs can drop another one to two orders of magnitude.

When costs drop this fast, quite a few usable scenarios can actually be deployed. Because I remember even a year ago or six months ago, much of the discussion focused on how GPT-4 was great but expensive.

But historically, technological progress is fast. Moore's Law drove computing chip power to become cheaper at a very rapid pace, and now large models are also dropping in price very quickly.

Many scenarios can now actually use GPT-4 level capabilities — things like customer service, internal enterprise productivity tools, communication with people. So in Silicon Valley, we're sensing that many enterprise service applications can now move from product prototypes to genuinely large-scale deployment. And in the future, GPT-4 level capabilities might become nearly free.

Like electricity — sure, you pay for it when you use it, but when you buy a computer you don't really think about how much the electricity costs. So I think this is something very important for AI to spread quickly.

Koji:

Have you seen any specific product that became feasible because of such a dramatic cost drop?

Yusen:

For example, previously many products had to do RAG, vector databases. But when your context window becomes much larger, and when both input and output token prices become very low, then maybe many things can just be put into context.

Because a company doesn't actually have that much text content — it can basically be called up at very low cost. So scenarios like customer service can be deployed very quickly.

Of course, we're seeing many products still in development, maybe not yet deployable in the short term or generating concrete revenue yet. But we can see an enormous number of companies working on this. So this is something that excites me.

Koji:

What else shocked you?

Yusen:

Second is the multimodal capability everyone's talking about. Especially what OpenAI demonstrated in their demo, which made everyone think — is this the scenario from the movie Her, where AI communicates emotionally with humans?

I think as human beings, we're social animals. So when we see AI able to communicate with us in a very socialized way, we do feel very shocked — that's certain.

For example, Google actually demoed a scenario where: a person filmed a video with their phone, and while chatting suddenly said, "Where did I put my glasses?" And Gemini told them: your glasses were placed in such-and-such place just now.

Many people's first reaction to seeing this was: so I can find anything I lose from now on.

I think this actually demonstrates the model's multimodal recognition capability. And also context — its very important role.

New Opportunities

Koji:

You just mentioned the movie Her. After the GPT-4o launch, this film seems to have returned to everyone's focus. When we watched it years ago, we probably thought it was science fiction, but today it seems to have become a documentary.

If people can really fall in love with AI, what kind of changes do you think it would bring to people?

Yusen:

First, I want to express a viewpoint: when we talk about AI and human emotions, we think of Her, we think of falling in love. In movies, these are very dramatic, very watchable scenarios. But actually, human relationships aren't all romantic — there's also friendship, family, trust, companionship, these all belong to emotions.

In human history, only another person could give a person emotional connection. Maybe cats and dogs, but basically only a person could give another person emotions, so it was a scarce thing. Think about it — in everyone's life, there aren't that many others who can provide emotional value.

Koji:

And there's actually another interesting saying: compared to needing love, people actually need trust and safety more.

Yusen:

Now, when this kind of emotional communication can be produced infinitely at low cost, it will bring many changes.

In human history, many changes happen when something originally scarce becomes no longer scarce.

For example, at first people all needed to eat food, but after food became abundant, people started pursuing higher spiritual enjoyment — though maybe also became obese.

If high-quality emotional value becomes less scarce, then people who originally needed companionship, who were depressed, who needed encouragement — various positive emotional needs might go from being unmet to being met, and the world might become a better place. But at the same time there will definitely be many negative effects. For example, telecom fraud at least used to require some guy behind the scenes tricking you — now AI might trick you in a very vivid way.

Yusen:

Previously, humans had never experienced another non-human being expressing emotions to them. So from this perspective, we're very susceptible — in English, very vulnerable. We're easily moved by an AI's emotions and expressions, because we've never encountered such a being before. Maybe ten years from now we'll all be used to talking naturally with AI, and everyone will be accustomed to AI's existence and won't be easily moved. But right now, whoever can build this technology first might have enormous impact.

Koji:

A friend recently told me: have you noticed, when we were kids watching Doraemon, in basically every episode Nobita would get bullied, run home crying, and rush into Doraemon's embrace. Doraemon loved him unconditionally, comforted him, figured out solutions for him. Nobody thought there was anything wrong with their hug — it felt completely natural: Doraemon could give Nobita comfort.

Actually human emotions are relatively easy to hack, and our needs can be satisfied by a Doraemon-like character. On this point, I'm actually relatively optimistic.

Koji:

From an investment perspective, how would you think about and judge the investment value of this recent wave of AI companionship-type projects?

Yusen:

I think AI that can communicate emotionally might be more like MSG — it makes many dishes taste better when added. But if you eat MSG straight, not to mention whether it tastes good, it might be poisonous.

For instance, every app we use today — whether hardware or software — communicates with us in a highly utilitarian way. Take your phone: it works well, it's fast, but you don't feel any emotional connection with it.

There have been some attempts. NIO, for example, has Nomi — a small display with facial expressions. But with current AI capabilities, or even AI capabilities from a year ago, this kind of interaction ends up feeling pretty artificial. It can handle very simple tasks, and mostly it just acts cute.

But imagine: if all the hardware you used could communicate with you in a human way — normal, emotionally attuned, warm — your relationship with many products would fundamentally change.

Yusen:

Imagine if your car could do more than just get you to your destination safely and quickly. It had seen many of the same landscapes you had. It knew where you often went. It had heard you speak countless times. Then, say you were driving yourself to the Great Wall, and it might say to you: "Three years ago when we came here, your child hadn't even been born yet. Now they're old enough to run errands for you." — Perhaps from that moment on, your understanding of this car would be completely different.

Later, when it came time to replace your car, switching brands would mean losing those memories. So you'd be more inclined to stay with the same brand, right? This transforms a product from having purely instrumental value to having instrumental plus emotional value. Emotional value can be a very deep moat.

Think about it: every company has a front desk person, right? They may not be the cheapest or the most efficient option. But often we don't easily replace them, because they understand the environment deeply, and we have a familiar rapport with them. So I started thinking: many products could develop emotional barriers, and many product interactions could shift from purely utilitarian to emotionally engaged.

Koji:

We're now supposedly in the "Hundred C Battle" (referring to the hundred-plus companies doing Character AI-style startups). What's your view on AI companionship projects?

Yusen:

Right now, many AI companionship projects are trying to build products that can fall in love with people. At present, I think this demand is relatively niche. Projects like Character AI, or domestic apps like STARFIELD's Talkie. Currently they seem to target a small subset of highly sensitive people who can easily empathize with AI, who really need that straight MSG.

But for ordinary people, it's probably more about adding a dash of MSG to scenarios we already have, making the experience better.

Koji:

I think the problem with Talkie or STARFIELD today is that the user has to initiate the conversation. And initiating conversation has a high barrier. So I don't think this is the ultimate form of AI companionship.

Yusen:

A quote comes to mind:

If you believe something is bound to happen eventually, do it every three years.

For example, the idea of an AI that can chat naturally with people goes back to Turing — that's why we have the Turing test. Humans have kept trying to build this, making incremental progress each time, never quite getting it perfect. But now we've finally reached a point where machines can communicate naturally with people. So which matters more — the form, or the content and context of the exchange? Right now we might find it remarkable that AI can talk to us, but soon we'll take it for granted, just as we now assume every screen is touchable. We'll simply expect AI to be able to converse. In that world, how do I redesign and upgrade my product?

A friend shared an interesting perspective with me recently that I found quite compelling:

Many of the companies that achieved massive success in mobile internet, whether hardware or software, essentially did the right thing visually.

Think about it: the iPhone was obviously about the great screen, the great camera; DJI is visual; TikTok is visual.

So in this AI era, the opportunity might lie in language and voice.

There are two things to get right here. One is deep language understanding — encompassing both natural language and programming language — and the ability to interact with people. For example, the programming agent applications people are building now: they're essentially about using programming language well. AI being able to code isn't just about writing code; it's about being able to start using tools. That could become enormous value.

The other is voice interaction. Voice or natural language interaction with computers has long been many people's dream, but it was genuinely cumbersome and offered poor experience. But going forward, take something like the Ray-Ban glasses that many people have been talking about recently — the Ray-Ban Meta collaboration works well because it's about voice interaction.

As language models get better, especially as on-device models improve and latency decreases, people may discover that voice interaction is incredibly natural. I think voice interaction, including device-side transformations, could present some very interesting opportunities.

Koji:

At a recent Microsoft event, one particularly memorable moment was: Copilot+PC has a dedicated hardware button — press it and you can start talking to your computer. And because the computer fully knows what you're looking at on your screen and what you've looked at before, the demo Microsoft showed was the computer playing games with me, telling me what to do next to advance.

Another user experience that recently moved me was Arc Search. As you know, to activate Siri you currently have to say "Hey Siri." But what Arc Search does: you open the app, lift it to your ear, and the app, sensing the motion trajectory and detecting the screen close to your face, immediately triggers: "Hi, what do you want to search?" Then you can start a conversation, consult with it like you're on a phone call. I believe this kind of experience innovation will emerge very rapidly across all sorts of places as technology evolves.

Yusen:

I very much agree with this point. Look at the GPT-4o demo — the whole thing was impressive. But one thing that looked rather odd was how they all had to hold up their phones to look around. Because of this, you realize that the phone we normally consider the most convenient device actually becomes an obstacle to the experience in this context.

Of course, we've also invested in some companies in the AR/VR space. But people keep asking: when will smart glasses truly take off?

I later thought: previously we were all thinking about the display side — we were trying to put screens on our faces. This led to two situations. One is the Apple Vision Pro scenario: we put a great screen on, but it's heavy. The other is earlier simple glasses: we put a not-so-great screen on, it's light, but not very good. So you either have a weight problem or a display quality problem. Putting screens on glasses, I think, remains difficult when there's no fundamental breakthrough in the underlying technology.

But glasses have another function: cameras. Previously, without AI, putting cameras on glasses for photography and video wasn't that important. I can always pull out my phone to take pictures — not much more trouble. But if we gradually develop a need to always have AI see what we're seeing, then putting cameras on glasses becomes much more meaningful.

Koji:

Reminds me of that MUJI song: seeing what you see, living the time you live.

Yusen:

Right. And I also wanted to ask you — I saw you posted something on Moments recently, saying you believe AI will exacerbate wealth and resource inequality in the world, but that AI can help people achieve emotional equality. Could you expand on that? I'm quite curious about your view.

Koji:

In the past, I believe emotions were massively unequal. We all know people around us who've been subjected to PUA, and friends currently struggling with depression, perhaps because their parents didn't have good communication skills when they were young.

What do I mean by emotional equality? In my definition, first, emotions are seen; second, emotions are processed appropriately.

I think in the past, many people's emotions were invisible. If you cried, parents would tell you to be good, don't cry — they couldn't see that your crying represented some underlying emotion.

But I have a very optimistic expectation: I believe AI can help humans at near-zero cost going forward. Your joys and sorrows can all be seen by AI, which can then interact with you in a relatively reasonable, warm, safe way. That's emotional equality.

Yusen:

Previously, high-quality emotional availability was scarce. Most people don't have followers, don't have suitors — they're just interacting with a few colleagues or family members, right? But AI can give everyone an excellent listener, and not just listening but giving good feedback. Previously, you might have paid a lot of money for that.

Koji:

Today's AI psychotherapy products are questioned by many: they can't do better than human psychotherapists, so cheap or not, they're unusable. But my own view is: AI psychotherapy doesn't need to compete with human psychotherapists on quality, but on time. Because no matter how good a human psychotherapist is, you only get one hour a week. But AI can provide 7x24 on-demand support.

A Year in Review

Koji:

Standing at the end of May 2024, do you remember what AI looked like at this time last year? What did you look like?

Yusen:

My broad judgments about AI haven't changed much. Scaling laws mean we'll have more powerful models. We'll need more compute, more data, more exceptional talent.

But the pace of AI development has actually exceeded what I imagined. Because look at this past year: we've had tremendous progress in multimodal understanding. And on the generation side, we've had Sora, we've had Suno.

In these domains that people previously considered very difficult, suddenly we've had their GPT-3 moments again. They've generalized to a certain degree, and quality has reached acceptable levels.

And on cost, we've seen more than tenfold changes. For something to drop tenfold in cost in a year, even several dozen fold — that's extremely difficult in any domain.

I had a judgment last year that the model arms race would keep going, but application deployment would be slower than people expected. But now, I actually think I need to revise my expectations for application deployment upward somewhat.

Because we're indeed seeing that as AI costs drop and its capabilities generalize, more and more people are trying to use AI.

Koji:

Over this past year, have there been any significant changes in your investment strategy and thinking?

Yusen:

I don't think there have been dramatic changes, but I've become clearer about one thing: in the early days of a technology, we should invest in young pioneers who have deep understanding of that technology.

For instance, we can divide historically technology-driven entrepreneurship into three stages.

The first stage is research-driven. At this point, whether you can build it at all is what matters most — it's a binary, you-have-it-or-you-don't distinction.

The second stage is engineering and product-driven. You can build it, I can build it, but maybe I build it better than you. That's the second stage.

The third stage is operations and business-driven. Everyone can build a decent product, but who does operations and commercialization well matters most.

When we actually entered the internet industry, we were already in the transition from product and technology-driven to business and operations-driven. Look at when we started our companies — we thought building an app still required some skill, right? So back then there were companies specifically helping people build apps. But later, building apps became trivial.

Eventually, it became about growth hacking. Growth became hard. Then it became about how everyone does user acquisition, how everyone does monetization.

Yusen:

In the past, say, if I wanted to build a robot nanny, I knew what I wanted to make but I couldn't build it. Now, we've returned to a stage where research matters enormously again. So the type of entrepreneurs we support should change. At the same time, our thinking about what's worth investing in has also changed.

The internet, in its later stages, was actually idea-driven. Because when your engineering and product capabilities become very strong, what you're missing is ideas. Like, I need to build a Pinduoduo, or I need to build some new social network — these were idea-driven. So idea-driven has two characteristics. First, everyone follows very quickly. In the later stages of the internet, whenever something emerged, suddenly within a year hundreds of companies would follow. Because as long as I saw it, building it wasn't hard.

Second, garage startups. Garage startups mean it's not that difficult, so as long as I have an idea, a few people in a garage can hack it together.

Now in AI, first, the ideas have been there for a very long time. The robot nanny I mentioned, Her, including these coding agents that help you program — these are things people have thought about for ages. It's not that they couldn't imagine them; it's that they couldn't build them. So that's the first point. Second, this means the teams you invest in must have deep understanding of the new technology inside this space, and they also need to accumulate resources — they genuinely need to make something that was previously very difficult actually achievable.

I often used this analogy: suppose we go back to when we all used BlackBerry or Nokia. Suppose someone time-traveled from the future and said, in the future there's an app called Douyin, it's incredibly awesome, you need to build Douyin now. We'd say, got it, but how do you build Douyin? How do you build Douyin on BlackBerry and Nokia? So back then, having the idea was useless — you needed the technology.

Now we're in a stage constrained by technology. So we're now focused on finding young, excellent technical entrepreneurs — this is actually what we're prioritizing in our search now.

Koji:

Speaking of which, I thought of an old friend I saw last week, hadn't seen him in over ten years. Back then he was doing part-time iPhone app development for Jiepang, but his main job was as an engineer at Apple. What problem did he help us solve? When we first launched our app, the screen would be black for three seconds on startup, because there was no concept of a launch screen back then. So we were probably the first in China to do a launch screen, so during those three seconds you could see a colorful image.

I was thinking: damn, there was actually a time in history when you needed to find a guru just to make a launch screen?

Yusen:

Because we've all lived through mobile internet in the past decade and become so-called veterans ourselves, we always want to think about endgames, right? I think thinking about endgames is a good thing. But when it's too early, forcing yourself to think about or answer what the endgame is — that's actually very hard too.

When we started our companies, how did we know what our endgame was? So I'm thinking: don't apply mature standards to young industries and young founders.

Koji:

If we say large models are now an elementary school student, we hope it can happily and carefreely learn knowledge, grow up mentally and physically healthy into a PhD as soon as possible — don't fleece it too early.

A year ago, on ZhenFund's podcast, you said one thing: application-layer companies need to build their own models in the long run, and model companies need to build their own applications in the long run. You just talked about three stages: technology, product, and operations/commercialization. Listening to this, it sounds like you think the first two stages should be done by one company? Has this view changed at this moment?

Yusen:

I think if you want to build a top-tier application-layer company, at least from a global perspective right now, you still need some control over models. It doesn't necessarily mean you need to train a foundational large model from scratch, but you need substantial control over continuous training of your model's data.

Companies that can really figure out models will have advantages when building applications.

But indeed in the long run, I think maybe not. In the long run, when these advanced AI models become very cheap and universally available, maybe everyone will increasingly build on the upper layers.

If we use the internet as analogy, in the early days many software or internet applications actually needed to do substantial work at the network communication layer or chip layer themselves.

But now when building an application, say, the languages you write in have become high-level languages, right? And your infrastructure has become cloud services — you don't need to think about how your server room is laid out.

Koji:

I even helped Xing Wang move servers to a data center near Beijing Railway Station back in the day. Thinking about it now feels like another lifetime.

Yusen:

Now your application can get huge without ever thinking about this. So the analogy is: now you need to think about building your own model, but maybe AI entrepreneurs ten years from now won't need to think about this at all.

Advice for Entrepreneurs

Koji:

Then for AI entrepreneurs today, there's another shared challenge: how do I choose my 0-to-1 entry angle. Especially recently, after ByteDance's launch event, their Volcano Engine website pulled out the full list of all their AI products. Dense like a contemporary hao123 — it gave me this feeling of, what do they call it, "kids make choices, adults want it all."

Under such an aggressive big-tech backdrop, how would you advise entrepreneurs to choose their 0-to-1 entry point?

Yusen:

Right now we can only give some relatively simple or general ideas. Because the whole thing is so early, we're still studying it too.

But generally speaking, a universal rule of entrepreneurship is to first find a small enough but sufficiently painful entry point, right? You'll actually find these big tech companies are all doing AI now, but there's a problem: because they take AI very seriously, they treat it as a big deal. So when some needs don't seem that large, they might not actually spend much time on them. I had a summary before: in the novel The Three-Body Problem, to protect Earth, the final strategy was to make aliens look at Earth and go "ugh." Because if they think Earth is a threat, they'll blow it up. So they came up with many schemes to make Earth look unattractive, or just "ugh."

I think in entrepreneurship, whether internet, mobile internet, or AI entrepreneurship, there's something similar: you can't let big tech look at what you're doing and immediately love it. Often, you need to find domains where big tech looks at them and goes "ugh."

So what makes big tech go "ugh"? For example, some things people think are too low-class — originally, things like Toutiao or Pinduoduo, at the start people more or less thought they were low-class, right? Some of these things that seem not so sophisticated, often big tech doesn't do.

Some are too exhausting — like earlier DiDi, Meituan, these had lots of ground promotion work, lots of big tech wasn't willing to do it, found it too exhausting.

Some are too niche — like Poizon, Bilibili, often big tech thought too niche, not worth doing.

And some previously had legal risks — like early DiDi, Uber, actually had lots of legal risks; and Airbnb, how can you sleep at someone else's house?

Koji:

When we say big tech, big tech isn't a simple individual — we're actually talking about the bosses inside big tech and the countless middle managers inside big tech. So many opportunities aren't that the bosses don't want to do them, but that the middle managers aren't willing to do them.

Yusen:

So I think often it comes down to finding a domain where, first, users have this need — they have an unsolved problem. Second, for big tech, there are some relatively simple reasons they haven't done it. For example, earlier with many generative models, why do American big tech companies like Google still do such poor generative models? Because they're constantly avoiding copyright challenges.

For example, if we ask Google people, they can't use YouTube data for training — it's quite funny. Everyone in the world wants to train on YouTube data, but Google can't. Because if Google wants to train on YouTube data, they have to go to the YouTube division. YouTube thinks: I'm the content producer, you use my training to generate video, then I have countless questions for you, right? And because it's Google, they actually can't skip over this.

Silicon Valley and China

Koji:

This year you've spent a lot of time in Silicon Valley, shuttling between China and the US. I also think today we seem to have returned to a period of observing Silicon Valley, learning from Silicon Valley. Yusen, what specific aspects of Silicon Valley today do you think Chinese entrepreneurs should learn from? But conversely, I also want to ask: what stories of Chinese entrepreneurs today could be reverse-exported to Silicon Valley?

Yusen:

I think the emergence of this whole AI wave actually originated from researchers like those at OpenAI. Whether it's scaling law, or innovation around the transformer architecture, or the current connectionism overturning the previous symbolism — the paradigm innovation in AI research — it still comes down to being the people who define what the next generation does.

I do think that because of the research environment, capital environment, and other factors, entrepreneurs in Silicon Valley are probably more willing to think big — to pursue what we might see as pretty crazy ideas. Maybe everyone wants to get a PhD instead of going straight to work after high school. I think this has a lot to do with their environment.

I've actually been thinking about this: entrepreneurship seems like taking an exam, right? We often have founders come to us and ask, "What are you investing in lately? What's hot right now?" The subtext is: whatever you're investing in, I'll do that. Or: you set the question, and I'll solve it for you. But I think the small number of truly exceptional people we're looking for are actually the ones setting the questions — they're creating the problems for the market to solve. Once they pose that question, then many people in the market start trying to answer it.

For example, Moonshot AI — they were very early on the long-context text thing, basically the first to make a name for themselves domestically. Of course many people followed afterward, but I think Zhilin Yang and his team had a strong technical vision here.

Although they weren't ahead of the market for that long, they were at least among the first to say: we're going to devote our limited resources and energy to long-context text, and build a good product that others will follow.

Yusen

Second, I think we should stay optimistic about the future. In Silicon Valley, there's also a lot of debate about whether AI is a bubble. I think it definitely is a bubble — it's just a question of how big. There are doubts about commercialization, about whether we need to go into this bigger model forest. But I think the overall atmosphere is still much more optimistic.

Of course we might have all sorts of rational reasons to be cautious, to be pessimistic. But I think if we've chosen entrepreneurship or VC, we should approach the future with an optimistic mindset. Because there's a saying:

Pessimists are often right, but optimists are often successful.

Or rather, successful people tend to be optimists. So maybe we have this innate optimistic bias — we still believe it will create tremendous value, that model capabilities will keep getting stronger and cheaper. There will be many opportunities in this state that make people more willing to take risks, to try things.

On the other hand, although a lot of AI progress first happens in Silicon Valley, we've invested in applications like HeyGen and Monica.im that are popular worldwide. Recently HeyGen also got investment from Benchmark — one of the best VCs in America — at a $500 million valuation. When I went to talk with Benchmark, they said this was the largest first check they've ever written in their history. I think this is also Silicon Valley's recognition of Chinese entrepreneurs and Chinese products.

I think this always shows that despite all the geopolitical noise right now, truly good products have no borders. Chinese entrepreneurs can build many excellent products, and gain recognition from both users and capital. Actually I think this is why people shouldn't be so pessimistic.

Because many people say, with all the geopolitical factors now, are we going to see Chinese AI and American AI diverge? But I want to say, in the end it comes down to whether we're creating value for users and customers. Because the world is still connected, right? If you create value, naturally people will want to use it, and naturally capital will want to invest. So I don't think we need to be so pessimistic.

Koji

Yes, so at Crossing we've recently distilled a slogan: finding and bringing together proactive actors in the AI era. Because without action, there will be no results.

Looking at deals on both sides — Silicon Valley and China — both seem super excited about AI right now. From the topics people care about to the concentrated directions of VC investments, have you noticed any obvious differences?

Yusen

I think what's very easy to see is that in America, if you look at the past decade or so, most of the big things have been in enterprise services, B2B. With AI, everyone can see it playing a role in many enterprise service scenarios — whether it's programming, customer service, internal workflows, or various sales scenarios. So there really are a lot of people in Silicon Valley doing AI entrepreneurship in the enterprise services direction.

Actually their thinking is very simple: I'm using AI to build better enterprise software. They used to worry that AI inference costs were too expensive, because enterprise software was very profitable — your marginal cost was zero, so your margins were high. They worried that if every use incurred inference costs, it would change their margin structure and make it less attractive. But now they've realized it's probably like electricity — it has a cost, but you can basically ignore it; it won't change your marginal profits that much. So they're actually increasingly focused on B2B application opportunities.

Yusen

Of course they invest in B2C too, it's just that B2C maybe isn't like what we've seen in recent years where B2C applications have advanced by leaps and bounds. And why is enterprise services so hard to land in China? That could be its own podcast episode — we could talk about it for a very long time.

Because the biggest application playbook in China over the past decade has been: build something that strongly attracts users with a good flow, capture their time, then sell them ads. Everyone's thinking about how to build the next ByteDance. So I joke that this is also how different domain experiences lead people to look for different things.

If you live in Shandong, you probably only want to take the civil service exam. But if you grew up in Silicon Valley, you probably want to start a company every day.

That's the impact of different environments. It is true that in America there's probably more conviction in scaling law. So as Sam Altman said, you should probably take continuous model upgrades as a given.

I often ask founders a question: if OpenAI holds a launch event, do you get excited? Because you can use better models to serve your customers. Or do you feel nervous? Because you might get disrupted by a more advanced model.

We hope to invest in companies that feel better when OpenAI or Gemini gets more awesome — not ones that worry about being disrupted by the models themselves.

Yusen

I think one problem China faces is really: I have the existing models, how can I generate revenue and survive? Because the capital environment may not be that great. Of course I think this is the current reality, but I also feel that capital markets are often like a manic-depressive patient. Sometimes they're extremely optimistic, sometimes extremely pessimistic.

For example, I remember just three years ago, many Chinese companies were valued higher than American companies. If you look at PE/PS ratios, now people won't even invest at 10x PE. I think being too optimistic isn't normal, and being too pessimistic isn't normal either.

Warren Buffett also said: Mr. Market is always a pendulum, swinging between optimism and pessimism, right? I think right now there's a lot of pessimism about future predictions, but I think it's worth considering: if new technology really brings a lot of synchronization, maybe our emotions will become optimistic again.

Koji

That's actually possible. I feel like America's enthusiasm or momentum for B2B investing is crushingly dominant over B2C.

Recently, when Crossing analyzed what 260 AI companies backed by Y Combinator actually do, seven or eight out of ten were B2B companies.

Yusen

Because America also hasn't had a big B2C thing in a long time.

Koji

The last one was probably Snapchat.

Yusen

That was ten years ago.

Koji

Right, so America's entire capital base, talent pool — none of it is there. This is actually an opportunity for Chinese entrepreneurs.

Yusen

Right, but look at what China has exported to the world in recent years — companies like TikTok and Shein. Maybe what each side is good at is just different.

The Future of VC

Koji

For Chinese entrepreneurs today, the primary market environment isn't optimistic. At the same time, considering today's large model capabilities and the optimization of development tools and platforms, building something from zero to one has become easier — the startup capital and team size needed seem to be smaller.

So it's also trendy now to be an indie hacker, an independent developer, or to build a so-called bootstrapped startup — not relying on fundraising, but making money first and rolling it up bit by bit.

Under this trend, how do you understand whether the value played by VC, especially angel investment institutions, has changed? Today's ZhenFund, ZhenFund a year ago, ZhenFund ten years ago — are there changes already happening?

Yusen

First, I don't quite think what you just described is the full picture.

It's true that AI capabilities let one person write more code, or help one person complete more automation. But if you want to build a good AI company, you probably need more data to train your model, and more compute. At the same time, you probably need better talent, better AI scientists, right? So the resources needed here, whether in terms of money or people, may actually be increasing.

So I think on one hand, yes, using AI to quickly build small applications has become very easy. But at the same time, making truly innovative, defensible breakthroughs in AI has actually become harder, right? So this probably still requires the support of investors and various resources.

Second, what I keep emphasizing is that AI is currently in the early stage of technology. The early stage of technology and the mature stage of technology are very different. In the early stage, trial and error is very important, because no one knows what the next direction is. At this time, high risk tolerance, angel investors willing to support wild ideas — this becomes very important.

Yusen

Actually you find that when an industry reaches maturity, first of all, many people in it are already very rich. For example, look at China's "New Energy Vehicle Trio," right? When these three founders started their companies, they were all already billionaires — they already had a lot of money. And second, maturity is often the opportunity for old hands, serial entrepreneurs. People like Zheng Huang, Yiming Zhang, Xing Wang — they're all very experienced people.

These days, we're seeing many founders — especially young ones without prior entrepreneurial experience, coming straight from academia — becoming increasingly important. Because these founders don't start out with significant accumulated capital, I feel that angel investing actually matters more in the early stages of a technology cycle.

Once a technology matures, perhaps industrial or strategic investment becomes more important. Because people increasingly know what to invest in and what to build, I actually think ZhenFund is well-positioned for this moment.

ZhenFund actually has the most AI unicorns in its historical portfolio. Many people think of us as investors in companies like Xiaohongshu, or Perfect Diary, or previously Jumei — these consumer projects. But in reality, AI projects account for more than half of all the unicorns we've invested in.

Later I started thinking about this: people working in AI are basically individuals we can relatively easily identify as exceptionally accomplished academically, because the barriers to entry are so high.

But at lunch today, I was eating with an investor in CHAGEE. CHAGEE is doing incredibly well, but the founder is genuinely quite grassroots — honestly, it would be very difficult for us to assess from scratch whether someone could build a tea company doing tens of billions in annual revenue. So I've realized that while we have some good examples in consumer investments, it's actually very hard to judge from the person alone. Since we can't invest at late stages, once a company has already taken off, we can't get in anymore. But with technical projects like AI, if we grab the most talented people, we can actually cover them quite well.

Koji

Because from Bob Xu onward, including yourself, ZhenFund has been very actively encouraging investors to focus on studying people. Everyone knows you invested in Zhilin Yang of Moonshot AI very early on. How did you and Zhilin first meet?

Yusen

We were paying very close attention to Zhilin Yang while he was still in school. We continuously cover Chinese talent at these top schools. We ask ourselves: who are the five smartest people at this school? Zhilin entered our radar very early — at Tsinghua he was known as the god among academic gods. Even a Tsinghua slacker like me understood what that meant, so we paid very close attention to people like him.

When Zhilin first started Recurrent AI as CTO, we actually invested in Recurrent AI. A very important reason was that Zhilin was one of the co-founders. So when ChatGPT came out, I was actually studying AI intensively. I was thinking: who among the people we've invested in are in this space?

I asked Zhilin: what do you think about the opportunities brought by large models? He said he was actually about to start a new company. So we talked then, and we became one of his first-round investors.

This also reflects how our strategy of focusing on people rather than things is especially helpful in the AI era.

Of course many people are studying AI now, and I'm studying it too — I read papers every day. But I can say very responsibly that AI is changing so fast, very few people can truly claim they understand it.

Look at how Turing Award winners are still arguing with each other every day. Just yesterday Yann LeCun was fighting with Elon Musk. And at the same time, Yann LeCun thinks Transformers are unreliable. So is Yann LeCun reliable? If Turing Award winners haven't reached consensus, how can a small VC like me make judgments?

So I think first of all, we shouldn't believe we can truly understand AI.

And I've noticed that two years ago, nobody predicted ChatGPT's emergence. In June 2022, I went to Silicon Valley and talked with Sequoia, Founders Fund, Benchmark, and many top US fund investors — not a single person mentioned having any sense or judgment about AI.

But three months later, Sequoia published that AIGC article. Then three months after that, ChatGPT came out, and everyone pivoted to AI.

So I think we should maintain an ordinary mindset toward frontier technology — it's hard to say I can understand it, and it's hard to predict what will happen in the future.

I often think: if ChatGPT only happened a year and a half ago, how can we claim to know how AI will develop two years from now?

So at this moment, we grab what doesn't change here: regardless of how things develop, the people building it must be those who deeply understand AI, who have done extensive research, and who are fighting on the front lines — this is the core logic of our people-focused investment approach.

Big Industry Questions

Koji

So with AI changing so fast, keeping yourself on the front lines becomes very important. Beyond continuous observation and thinking, I think another crucial aspect of being on the front lines is asking good questions. Among the questions people are generally curious about today, which ones do you think are particularly worth asking?

Yusen

I think people are actually concerned about many similar things. For example: where is the ceiling of this wave of large models or existing technology? What are the constraints in reaching this ceiling, and at what growth rate should we set expectations? Of course, as I mentioned, even the top AI researchers and big shots have many disagreements on these questions. So I think we're constantly adjusting our probabilities of future outcomes based on new inputs.

It's somewhat like a Bayesian model. In a Bayesian world, before ChatGPT, AI was in an unactivated state. But when ChatGPT appeared, you immediately found that the probabilities of many future things changed. So we should think about this quite flexibly.

Second, I think it's about what scenarios AI-native applications will emerge in.

Here's how I think about this: when we look at the internet, mobile internet, or earlier semiconductors, new technologies typically solve old problems first. For example, the internet for sending mail was called email, the internet for reading news was called web portals, the internet for moving offline stores online was called e-commerce.

But when internet penetration reached a certain level, things changed:

For example, once we were all online, social networks emerged — we met on Xiaonei, something that didn't exist without the internet. This is what you call a native model.

For example, once information was all online, search engines were needed — also a native model, giving birth to Google.

For example, once buyers, sellers, merchants, and money were all online, platform e-commerce emerged — another previously nonexistent model.

And these models are where the real trillion-dollar companies come from.

So when mobile internet began, people initially looked for mobile search engines and found Baidu and Google; looked for mobile browsers and found Chrome; looked for mobile YouTube and found YouTube.

But later, as mobile penetration increased, things changed again:

For example, once content creators and producers all had smartphones and 4G, Douyin, Kuaishou, and Xiaohongshu were born.

For example, once blue-collar workers all had smartphones, DiDi and Meituan emerged.

For example, once gamers all had phones, miHoYo became possible.

So you see again: when a new technology's penetration reaches a certain level, new so-called native business models are born, creating enormous value. And these companies are all startups.

So I think, first, there's no need to rush. Right now, the priority should be driving up penetration. We should first invest in companies or products that get more people using AI, right? Investing in compute, investing in data, investing in leading models — these are all aimed at achieving this, at getting AI penetration up.

Once penetration rises, native business models that didn't exist before will gradually emerge because of that increased penetration. For example, if each of us has several AI assistants handling different things for us, might there be entirely new business models targeting these assistants, or the very fact of having them? Or if most content in the world is AI-generated rather than human-generated, might content generation platforms, creation tools, and consumption platforms undergo significant changes?

Koji

This is also what everyone cares about most: what are the native applications of the AI era?

Yusen

I think although this is what everyone cares about most, this question can be set aside for now. Everyone should first drive up penetration, first find usage scenarios, and gradually, people will naturally discover what the native scenarios are.

So I have this question too, but I don't necessarily expect an answer immediately. Or rather, I think the process of searching for such an answer is itself a very good process, a very important one.

Koji

All the big questions you mentioned above — whether startups or Silicon Valley giants — everyone is enthusiastically trying to answer them. Among the Mega 7 companies and their CEOs in Silicon Valley, have you seen anyone who impressed you?

Yusen

If you're talking about Mega 7 CEOs, I think much of it comes down to how they can break out of the organizational rigidity trap of big companies to get things done.

First, many people mention Satya Nadella. As Microsoft's CEO, he truly transformed and refreshed the entire company — just like the title of his book.

What impressed me most over the past year was during the OpenAI palace drama, how incredibly fast he moved — the same day or the next day — to express support for Sam Altman, to express willingness to acquire him. This speed is absolutely super fast for a multi-trillion-dollar company. You can imagine how many people he had to convince in that process — board members, business line heads — and how strong a conviction he and Sam Altman must have built to do this.

But this decision wasn't made in a vacuum. His prior investment in and support for OpenAI, his conviction that this team could produce ChatGPT, and his series of earlier acquisitions like GitHub that deeply transformed Microsoft from one of the most closed-source companies into one of the broadest open-source knowledge companies — these were all remarkably impressive strategic moves.

Also Elon. Because this time in Silicon Valley, I use Tesla FSD almost every day, and I've genuinely found that as FSD has gotten so good, I increasingly don't want to drive myself. And I heard from them that FSD's visual end-to-end is actually being built by a very small team. And after ChatGPT came out, both the team and Elon saw the capabilities that end-to-end models and large models could bring.

Because legacy autonomous driving was modular — different modules governed by different rules. Elon tore it all down and started over, abandoning tens of millions of lines of code to train a pure vision-based end-to-end model from scratch, reaching and surpassing the previous FSD's capabilities within a year.

Later I kept thinking: how do you justify a decision like that? Saying you're going to discard a decade of work, that everything from the past ten years was for nothing, and starting fresh — without even knowing if you can pull it off. I think this is probably something only a certain kind of founder-CEO could do.

We've all been inside companies. To tell your team that so many people, so much money, ten years of work — all of it, worthless, scrap it — that's genuinely incredibly difficult.

Koji

What about Sam Altman?

Yusen

Honestly, I don't think any of us are really in a position to evaluate Sam Altman. There's too much information we don't have, right? And too much gossip and rumor swirling around.

But I think there's one fact we can observe. OpenAI used to be a research lab working on several research themes. DeepMind was a research lab. Boston Dynamics was a research lab. Yet the vast majority — almost all — of these research-lab-style startups ultimately failed to find their commercial product and path. DeepMind and Boston Dynamics both got acquired, right? Acquisition became their exit. Because they had so much talent, people said, "Your talent is too valuable, we have to acquire you."

In 2019, before Sam Altman became CEO, OpenAI was a nonprofit research lab running multiple research streams with no commercial applications. Then Sam arrived — and I'm not saying there's causation here, just a temporal relationship — and a few years later, OpenAI became a company that built an extraordinary product, with both B2B and B2C commercial models. They may not be profitable yet, but they're generating substantial revenue.

What happened? How much of this transformation is attributable to Sam Altman, and how did he accomplish it? I think this is extraordinarily difficult. Historically, perhaps only Google really managed this — going from a research project to a company with killer products to successful commercialization.

As for what role Sam Altman played in OpenAI's transformation? Maybe in five or ten years we'll know clearly from a book or something. But for now, we can see this change, this incredibly difficult change, and I find that quite remarkable.

Koji

This past year you've met AI founders — and we've been talking about giant CEOs. Back to AI founders: how do you see them as similar to and different from the mobile internet era's entrepreneurs?

Yusen

I think good AI founders today first need deep technical understanding — this may differ somewhat from the mobile internet era. Back then, some founders didn't really understand technology; they just needed to find people who did and have them build it, right? I think AI founders now probably do need genuine, fairly deep understanding of AI's evolution, including how to train and use models. This is also why, I believe, someone like Zhilin Yang stands out — having that real depth of understanding matters.

As I keep saying, we're still in a technology-constrained period, not an ideas-constrained period. Right now, many things are like, "This is pretty good, but you can't quite articulate it, right?" So how do you know what to build with today's technology? This is what finding technology-market fit becomes about.

Koji

You need to know the technology boundaries extremely well.

Yusen

Right, that's crucial. Second, I think in the later mobile internet era, we had a lot of very China-specific innovations, right? Back then I was actually traveling to Silicon Valley less frequently. I felt China had so many unique innovations, and such a unique market, that Silicon Valley playbooks often didn't apply here.

But now I think AI has arrived at another technological revolution, and North America — Silicon Valley specifically — is the first mover, right? They have the best people, the most cutting-edge understanding, and the most compute. So from both investor and founder perspectives, being internationalized can bring cognitive advantages. I remember we used to joke that Xing Wang was the fastest person at copying American apps for China, right? Of course, Xing Wang eventually built Meituan, which then got copied back to the United States.

Having frontier understanding matters at times like this.

But I think founders across all domains — whether selling milk tea or building AI — share a lot in common: strong learning ability; strong leadership to build excellent teams; creativity to think of what others haven't; and tremendous willpower to persist. These are the common threads.

Koji

What's your prediction for GPT-5?

Yusen

Talking with various researchers from OpenAI and DeepMind in Silicon Valley, the general consensus is that improvement in reasoning capability matters most — that's the outcome.

Because now we're finding the pace of ceiling-raising somewhat uncertain, since GPT-5 hasn't launched, so you don't know how high the ceiling's been pushed. But deployment below is happening fast. As costs drop, how much the ceiling can rise becomes critical.

And I think the ceiling's most important manifestation is this reasoning capability improvement. To solve reasoning, it comes down to data and compute, right? From those two angles, the data bottleneck now is first, how to generate high-quality synthetic data. Second, how to combine multimodal data with text data. Because people are discovering whether multimodal data's addition helps, hurts, or is neutral to text capabilities — these are still very open questions.

Then on compute: what people generally know is OpenAI's largest cluster is probably 30,000 to 50,000 GPUs. And NVIDIA recently disclosed in earnings that they're building a 100,000-GPU cluster. And I heard this might be for X AI.

So how does compute reach the next scale? Because with these two ingredients — data and compute — how does each find another 10x of headroom? This is probably what brings the next generation of models.

Then too, there may need to be architectural breakthroughs — what comes after transformer? This is also a frontier research topic.

The First AI Art Festival

Koji

The Fair and ZhenFund jointly launched the first AI Art Festival. This actually originated from a meal Yusen and I had around New Year's. He half-jokingly, half-seriously suggested to me: The Fair's slogan is "We will eventually change the direction of the tide."

Now if you don't actively embrace AI, you might not be able to change the tide's direction — you might get changed by the tide instead. I remember that day, I went back to the office and initiated the project, and the first AI Art Festival was born.

Yusen

AI will bring enormous impact to our world. What's quite interesting is that we used to think AI would mainly replace STEM-type work — supposedly rigid, rote stuff. But actually AI may first replace liberal arts and art students, right?

We originally thought AI would first replace physical laborers, but it turns out it may first replace knowledge workers.

So we felt art was the domain that most embodied human soul, human emotion, what makes us human. But unexpectedly, once Midjourney came out, people discovered 99% of so-called painters and designers became seemingly indistinguishable from AI. The shock this brought has been quite significant.

Historically, painting first aimed at realism. After photography was invented, people said realistic painting was useless. Then various Impressionisms, Fauvisms, new painting forms emerged. So I think this is also an opportunity of destruction and creation.

Koji

For this festival we set the theme as Love, Hate, Passion, Enmity — because these four words have been humanity's creative themes since ancient times. We invited people to use AI to create works about love, hate, passion, and enmity, to see what fresh perspectives might emerge. We received over 2,000 AI-created works on this theme. As a judge reviewing these works, what were your impressions?

Yusen

First, I think the production quality was quite high.

This partly comes from AI model progress itself — a cat drawn by AI two years ago versus now, completely different quality.

At the same time, people are increasingly proficient with AI software. I saw some fairly complex prompt techniques, including combinations of multiple applications, so it felt very rich and diverse.

And what could improve? I think much of it still has a somewhat "AI taste." Perhaps because many current visual models have somewhat similar underlying datasets, so it's relatively easy to identify an "AI flavor."

Perhaps next is how to combine AI with artistic creation techniques. Because now I see many works that are one prompt, one output, but I think often artists remix existing techniques — how to combine with AI techniques to produce something neither originally creatable nor purely AI. This is a more end-to-end, more integrated and comprehensive art.

Koji

Alright, our final question today, also a small question I'm personally most curious about: Yusen has always been a heavy science fiction reader, but in most science fiction describing futures where humans and AI coexist, the perspective is pessimistic. Today in our podcast, Yusen also said: "Pessimists are often right, but only optimists have a chance at success." Could you please end by recommending a science fiction novel that optimistically views AI's impact on humanity — for everyone facing the generally low social mood of 2024?

Yusen

This question is actually quite interesting, because when I got it, I actually went to ask Perplexity. Because first, I couldn't immediately think of which science fiction novel has AI as positive — it's always AI destroying humanity, or having some conflict with humans, because that creates dramatic tension. So I asked Perplexity to recommend a science fiction novel where AI is positive. It gave me some answers, but I felt they somewhat missed the point.

But right in the middle of recording this podcast, it suddenly hit me — there's actually a perfect answer: Doraemon. Doraemon is a manga, sure, but it really captures that beautiful imagination people have about technology and AI. First, tech and AI have all this practical utility — they help you solve so many problems. But second, there's also this strong emotional companionship to it. It's not just a magic pocket; it's a robot cat. So I think great AI might be exactly this — combining serious technology with genuine human warmth. That's the kind of future I hope AI development eventually reaches.

So if anyone ever asks you, "What's AI good for?" — just think about what Doraemon meant to Nobita.

"AI Frontline Investors" Podcast Series

At Crossing, we interview people actively investing in AI today — those immersed daily in AI products and business logic — to serve as our guides through the unfolding stories of AI investment and industry analysis from an investor's perspective.

Subscribe to the Crossing Podcast

We track the industry shifts and entrepreneurial opportunities emerging from the new wave of AI technology. "Crossing" was Steve Jobs's metaphor for Apple — standing at the intersection of technology and liberal arts, where great products are born. As AI transforms every industry, we seek out, interview, and bring together the "active participants" of the AI era to explore and embrace new changes and new possibilities.

👦🏻 Host Koji: Co-founder of The Fair and Tangdao, and an active participant diving into AI himself. [Koji's Jike[3], A Self-Introduction from Koji[4]]

👧🏻 Host Ronghui: Works at a tech VC, former Silicon Valley correspondent for CBNweekly. [Ronghui's Jike[5]]

Join the Crossing membership community:

☀️ Daily AI Fresh News Brief

👫🏻 Encouraging dating, friendship, and finding future collaborators

🦀 Add our assistant on WeChat to join: Rwkfbcianvd

References

[1] 此话当真: https://www.xiaoyuzhoufm.com/podcast/646f194853a5e5ea1408d97c

[2] 十字路口: https://www.xiaoyuzhoufm.com/podcast/60502e253c92d4f62c2a9577

[3] Koji 的即刻: https://okjk.co/0JSUes

[4] 一份 Koji 的自我介绍: https://www.notion.so/About-Koji-415843ab8db74235b98f8b1f67da1930?pvs=21

[5] Ronghui 的即刻: https://okjk.co/0cbnYV