2025 Kickoff Conversation: The Critical Year for AI, The Dawn of the Agent Era | A Conversation with ZhenFund's Yusen Dai

The first week of 2025 marks a special crossover episode between ZhenFund's podcast "**Genuine Talk[1]**" and "**Crossing[2]**." We've invited our old friend Yusen Dai, managing partner at ZhenFund, to join us. **We not only look back at AI's rapid evolution over the past year, but also ahead at the major opportunities awaiting AI entrepreneurs in 2025.**

The first week of 2025 marks a special crossover episode between ZhenFund's podcast "This Is Serious[1]" and "Crossing[2]". We're joined by our old friend Yusen Dai, Managing Partner at ZhenFund. We not only look back at AI's rapid evolution over the past year, but also look ahead to the major opportunities for AI entrepreneurship in 2025.

Standing at the start of 2025, both Yusen and I feel incredibly excited—we believe we're witnessing a pivotal moment in tech history. That excitement stems from two major events: the launch of Devin, and OpenAI's release of o3.

That's why we're greeting 2025 with optimism, convinced this will be a year full of promise.

Six months ago on "Crossing," Yusen offered an analogy: large language models were still elementary school students—don't rush to put them to work. Now, with the release of Devin, a genuinely usable Agent product, Yusen believes AI programming has completed a crucial evolution—from "I ask, you answer," to "I ask, you write," to "I ask, you do." This breakthrough represents not just significant progress in AI programming, but also signals a wave of promising vertical AI entrepreneurship opportunities.

💚 Happy New Year—may you have love and hope.

Listen on WeChat:

Listen on Xiaoyuzhou:

2024 in Review: AI Technology Explodes, Model Progress Exceeds Expectations, Application Growth Significant

🚥 Koji

Let's start with my first question for Yusen: What's your overall impression of 2024?

👦🏻 Yusen Dai

I'm very happy to collaborate with Koji again, to share our thoughts on AI development and investing, and to have this chance to exchange ideas with everyone.

Looking back at 2024, if I had to sum it up in one word: fast. Because we've seen rapid iteration in both AI models and products.

I remember at the start of 2024, the most advanced model was GPT-4. There was a new benchmark called SWE-bench[3]—it took common task types from GitHub and had AI attempt to complete them. At that time, the state-of-the-art model GPT-4 scored 2.8 out of 100. By the end of 2024, the Sonnet 3.5 that people could actually use was scoring 50—solving half the tasks. And the just-released o3 scored 71.7 in preliminary evaluations.

Optimistically speaking, at this pace, within a year—that is, by 2025—we could see AI solving the vast majority of tasks on GitHub. This means that while not entire jobs, many individual tasks currently done by programmers can indeed be solved. At the start of 2024, ChatGPT still struggled with basic arithmetic—people would test it on three-digit by three-digit multiplication and it would often get it wrong. But now it's handling IMO-level problems with ease, and even on Frontier Math, a test set that's difficult even for mathematicians, o3 scored 25. This was endorsed by Terence Tao, who said the easier problems are IMO-level while the hard ones are frontier research-level—and now AI can do reasonably well on them.

🚥 Koji

What impact has this had on applications?

👦🏻 Yusen Dai

We invested in Moonshot AI, whose product launched on October 9, 2023—just before 2024 began. By the end of 2024, it had 40 million monthly active users. Considering it's a new app only about a year old, that user growth is remarkably fast.

I also remember during the February 2024 Spring Festival holiday, watching Sora's launch demo and feeling completely blown away. At the time, I wondered how long it would take and at what cost we'd get to use such a video generation model. But by the end of 2024, people were already using it. Products like Keling AI, Hunyuan, and others—including Google's Veo 2—were arguably better video generation models than Sora was then, and they were free, making people feel like it was no big deal. So people's standards for AI products have risen quickly—what amazed them a year ago now feels ordinary. But we always feel there's more to be done, much that hasn't landed yet, and the progress is genuinely rapid.

I also think many predictions and opinions got proven wrong. I remember at the start of 2024, if you asked Chinese investors and entrepreneurs, many wanted to build "China's Character AI"—everyone thought this seemed like a consumer application with lots of users, and there was talk of a "hundred-C battle."

🚥 Koji

Early in the year, many predicted that a defining trend of 2024 would be the "hundred-C battle."

👦🏻 Yusen Dai

I didn't predict that, but many did. In August 2024, Character AI announced its acquisition by Google, and people realized breaking out of niche audiences wasn't so easy. I remember in March, Cognition—the company behind Devin—released a demo video. Back then, nobody believed it; people thought the company was full of it, some even called them scammers, and there were debunking videos. Then in December, Devin launched, and everyone was shocked to discover it was real, that it could actually do many AI functions. We'll discuss this more shortly—this was another major reversal.

I also remember the OpenAI board drama at the end of 2023, when the entire OpenAI staff collectively voiced support for Sam Altman on Twitter, with the message "OpenAI is nothing without its people"—it was everywhere. But by the end of 2024, who knows how many people had left. In the end, even Alec Radford, a core researcher and OpenAI veteran, departed. Basically most of the early employees left. And at the start of the year, everyone thought GPT-5 was coming soon, but GPT-4.5 never materialized by year-end. What came instead was a different path—o1, o3, this inference-time scaling route.

So I think the year was full of changes, both rapid ones and many unexpected ones. This is probably just the normal state of an industry in its early stages of transformation.

🚥 Koji

Six months ago on that Crossing episode, you had a core view: "Large language models are still elementary school students—don't rush to put them to work and earn money, give them more patience."

When you said that, it implied that while technology was advancing quickly, commercialization was still far off, and mass consumer applications were still distant. Do you still believe that today? Or do you think the pace of evolution has exceeded what you understood then?

👦🏻 Yusen Dai

First, that statement had a specific context—people were asking "We've spent so much money training models, when will we earn it back?" In discussing the investment return cycle for model training, I think this fits the pattern of every tech revolution:

First invest in infrastructure and research, then products gradually find their footing in real-world scenarios, and finally commercial revenue emerges.

So looking back at the year, in specific domains where model capabilities excel—like coding—large models have indeed crossed the threshold where they can "work." As I mentioned earlier, on SWE-bench, solving only 2% of problems at the start of the year clearly wasn't job-ready, but now solving 50% is. Especially after ChatGPT 3.5 appeared, we saw products like Cursor, Windsurf, and Devin emerge—they can genuinely help programmers solve many problems and deliver substantial productivity gains.

From a revenue perspective, some native AI applications have grown rapidly after finding product-market fit. For example, Cursor is now approaching $100 million in annual recurring revenue (ARR). Another AI coding company targeting non-technical users, bolt.new, reached $4 million ARR in four weeks and $20 million ARR in two months—the fastest growth ever for an enterprise application. And a Stockholm-based company called lovable hit $4 million in annualized revenue within four weeks.

Including our portfolio companies: Heygen hit $1 million ARR in mid-2023, then grew dozens of times over 18 months to nearly $50 million ARR by end of 2024. Our investment Monica has also surpassed eight figures in ARR—all achieved in just over a dozen months. Whether overseas startups or our own portfolio, significant progress has been made in user growth. Like the Moonshot AI I mentioned earlier, now with 40 million users.

So I believe AI already has "work" capabilities in certain domains, but overall revenue still falls far short of costs. We need to remain patient—after all, ChatGPT has only been around for two years. We're still in a phase where model capabilities keep improving and unlocking new application scenarios. Only after application scenarios generate enough value can commercialization gradually unfold.

🚥 Koji

Actually, I think the speed of this technology diffusion is incredibly fast. The Cursor, bolt.new, Heygen, and Monica I just mentioned—except for Monica, because Red Xiao gave me a VIP membership, I'm a paying user of the other three. The spread of these technologies feels faster than the last wave to me. Even without network effects, there's a passionate group of explorers at the frontier constantly trying new things and enthusiastically spreading the word. Crossing is part of this too—Yusen and I share anything exciting we come across the moment we find it.

I have a strong feeling, and it's why we're recording this episode: I hope people won't just watch from the sidelines, treating new releases as mere version numbers that don't affect them. I really want everyone to jump into the wave, download these apps, experience them early, start using them early.

Yusen Dai

There's a quote I love from sci-fi writer Gibson: "The future is already here—it's just not evenly distributed." If you're just using a simple chatbot in your daily life, or if you haven't really used AI products much at all, you might think this is all just headlines—who beat the benchmark, who did what.

But in certain domains, like programming or digital art creation, I believe AI tools have already become indispensable to their work. So I've always felt that spending a little time or a little money to experience the latest AI products is well worth it. It's a great way to viscerally feel the progress we're making in certain areas, and a good way to see the future.

AI Technology Diffusion: How to Let Everyone Create, Not Just Use

Koji

And regarding the massive advances I mentioned for both digital art creators and programmers, I think their significance goes far beyond just helping those two groups. More importantly, they're helping ordinary people do the kind of creation that only programmers and artists could do before—that's the bigger meaning.

So don't think "I'm not a programmer" or "I'm not a digital art creator, this has nothing to do with me." What I'm saying is, this actually has everything to do with you, because you can now do things that only they could do before.

Coming back to this—Yusen, how many AI application startups did you talk to at ZhenFund last year? What's your overall impression? Do you feel like the pace of AI application deployment is accelerating?

Yusen Dai

Our team collectively probably talked to over a thousand AI application startups. Looking at my own numbers, I chatted with over a hundred, close to 200 founders. We do feel that with technological progress, the pace of AI application deployment is speeding up.

Specifically, I think three developments matter a lot:

First is model reasoning capability, including releases like GPT-4o and o1. As models' reasoning gets stronger, hallucinations decrease, so they can plan and complete more complex tasks.

Second is improved model programming capability. In the digital world, many tasks can be accomplished by writing programs. As we mentioned earlier, programming capability is growing very fast. When these common tasks can be solved through programming, at least in the programming domain and other domains that can be generalized as programming, task execution ability improves dramatically.

Third is computer use, first proposed by Anthropic—AI's ability to use existing software, starting from browsers to other applications. All the software humanity has built can be used by AI to solve tasks. So combining these, I think AI's ability to complete tasks has improved significantly.

I think Devin's release in 2025 was important because it was the first product to turn agent from imagination and prototype into real-world deployment.

I think we'll quickly see agent attempts across various domains in '25. Of course many will still be in relatively early stages, but I think a lot of interesting thinking will be put into practice.

Koji

So we'll spend considerable time talking about Devin, and about our expectations for agent-represented AI development next year.

Yusen Dai

We see that AI application startup directions in the US and China are quite different. In China, because enterprise services are still somewhat difficult to deploy, many entrepreneurs want to do To C applications. And within To C applications, many tend toward time-killing apps—various emotional companionship, AI chat variants. In the US, what we see is that across various verticals, people are trying to replace portions of human work, making work more cost-efficient. This is a major contrast between Chinese and American startup directions.

Of course domestically there's also a major direction where robotics is particularly hot—the entire embodied intelligence space has many new companies emerging, raising lots of funding, even to the point where we feel it's somewhat overheated. But overall, I think everyone is still very excited. Especially for young entrepreneurs—because before, people might have felt the internet era was nearly over, we post-80s were the beneficiaries of internet dividends, but what could post-00s do? Before AI took off, they felt there really wasn't much to do in the internet space. But now AI has shown them many new opportunities, opportunities belonging to their generation of young entrepreneurs. So as a fund that always focuses on young people, we still see many interesting entrepreneurs emerging, interesting projects emerging.

Koji

Speaking of this wave of entrepreneurs, what typical commonalities do you see in them? Beyond being more friendly to the young?

Yusen Dai

Being young is an inevitable characteristic as times progress. I think first, they generally have more international vision—information spreads faster and faster. In the internet era, when an overseas app became popular, China might take three to six months to produce a copycat. Now basically when something new appears overseas, there are news reports the same day, many summarized and translated by AI. So everyone is generally very aware of overseas model application progress.

At the same time, because the products being made are often international too—since going global is now a major theme. Models themselves have strong multilingual capabilities, so people often start with global products from day one. This was harder to see in the internet era, when people would typically say "I'll just make a product for the Chinese market." Now people start with two paths simultaneously, both domestic and international. I also see many entrepreneurs and teams being more AI Native—quite a few have experience in AI research or engineering practice, which is why they can spot opportunities earlier and implement them.

But at the same time, I think for young entrepreneurs, because they may not have experienced many internet business processes, they have some lessons to learn in areas like promotion and commercialization. At this point, some veterans—like teams such as Monica that we've invested in, who have been through much internet-era growth—do have advantages in this experience. But I think these are all learnable, and can also be improved through hiring and team supplementation, so we're still very confident long-term. We believe the new generation of AI Native entrepreneurs can build very interesting products, and can also make up for the lessons they need to learn.

Koji

Next let's talk about the changes in understanding of AI technological breakthroughs, industry shifts, and entrepreneurial opportunities from last year to this year. First I want to ask: what are some views you quite agreed with a year ago that you no longer agree with now?

Yusen Dai

Too many—so later I didn't really want to record podcasts anymore, because every time I'd speak I'd get slapped in the face.

But to do early-stage investing, especially looking at early technology, getting slapped in the face is the norm. Only by not fearing getting slapped in the face can you continue learning and growing.

A little over a year ago, what everyone emphasized was Pre-training. People were talking about how many GPUs you need, how big your cluster needs to be—this is also why NVIDIA's stock exploded. Because people simply understood it as: more GPUs, more compute, throw more data in, and good models come out.

Looking at late 2024 to early 2025, in Pre-training, from OpenAI and various industry-leading teams, there has indeed been a relative bottleneck.

If we say Pre-training is compression of intelligence, then the intelligence easily compressible in forms like text has been pretty much compressed now.

Ilya said in a talk, "All this text on the internet is like fossil fuel, the text humanity has accumulated over so many years, and now we've trained it all into our models. Next we need new knowledge—whether knowledge still in our brains that hasn't been extracted, or new knowledge produced through AI—and the growth rate of such knowledge isn't that fast." So I think Pre-training's "brute force for miracles" is something everyone realized this year needs to change.

A year ago I did also talk about some agent content. At the time I felt that with large models still having many hallucinations, the landing time for this kind of autonomous agent or L4-level agent would need to be relatively long. But currently, model reasoning capability, code generation capability, and tool use capability have indeed progressed very fast. This means in the digital world, if we look at tasks with relatively certain target outcomes, like programming, agent deployment speed has indeed accelerated a lot. We've already seen products like Devin that are no longer just an idea, but have become reality.

There are two key points here: one is how to better plan tasks, the ability to do longer-term, long time horizon tasks has become very strong; two is using tools, including writing code to use them and using existing tools. When both these capabilities become strong, agent deployment speed may be faster than people thought, especially in the digital world.

Third point is, a year ago people generally believed model size would keep getting bigger—previously people said maybe 7B, 70B, maybe up to 700B. But currently, the size increase of advanced models actually doesn't need to be that fast.

We can get increasingly better results with a 70B model, while also being able to run the same capabilities on even smaller models.

So in reality, these truly massive models will probably be used mainly for model alignment, or as teacher models. This is actually somewhat reminiscent of the early personal computer era. Everyone initially thought CPU clock speeds would keep climbing higher and higher, but after reaching around 3GHz, single-core frequency growth largely stalled. Instead, performance gains came from better architecture and lower power consumption. Like the human brain, it's not about getting bigger — it's about acquiring more knowledge and skills within the same size, becoming smarter. On this front, I think model cost reduction has exceeded expectations. While everyone knew costs would keep falling, we're now seeing the same model or same level of intelligence drop to one-tenth of its previous cost annually. This will unlock many application opportunities — these were things people either didn't clearly recognize at the beginning of 2024, or views that shifted during the course of the year.

🚥 Koji

One more question about shifting perceptions: Were there things you considered worth watching at the start of 2024 that didn't seem that important, but have since become especially significant?

👦🏻 Yusen Dai

First, I think as investors, our understanding of frontier research often lags somewhat. Certain things may have already reached consensus within researcher communities, while we're still catching up.

A major focus in 2024 was the rise of Reinforcement Learning (RL). As mentioned earlier, Pre-training has hit a bottleneck, and using RL during Post-training to continuously strengthen model capabilities — especially after the release of o1 and o3 — showed everyone that the Reinforcement Learning path still has a long way to go, and model capabilities can improve substantially. At the beginning of 2024, this was really only discussed within fairly small circles, and hadn't become common consensus even within the research community, let alone beyond it. So we've found that predicting the technical trajectory of large models or AI is always extremely difficult. RL talent is actually quite scarce, so everyone is building teams and making technical investments in this area.

Alongside this, a very important new scaling law emerged: inference scaling law — how to extend reasoning time to get better results. This was a major development last year, manifesting not just in model design but in how we design products. Because most current products, whether ChatGPT, Claude, or something like Cursor, involve real-time interaction with humans — I say something, it responds. So how do we enable it to spend more time at each step, or even let it autonomously plan and use tools to keep working continuously, without requiring my ongoing input? This "System 2" way of thinking — not speaking off the cuff, but deliberating carefully to arrive at better results. How to achieve better performance on this front, I think, will be very important this year.

Another thing that people didn't consider very important a year or two ago, but now seems extremely significant: we already have a lot of intelligence in models, but models previously had very little context. For example, when I ask ChatGPT a question, it really only has my input as its context. In fact, any smart person would find it difficult to answer a question with just one sentence to go on. But now we're seeing, for instance, that Cursor can take an entire organization's codebase as context — you can select a large block of code as its context. And Devin is essentially integrated into Slack, able to draw on existing conversation records and feature documentation within the organization as context. When a model has more context at the same intelligence level, it can better understand intent and better answer questions.

On this front, new product designs that let users painlessly and easily bring in more context will become important. So the current ChatGPT-style back-and-forth, I think, is still a very primitive format. Everyone is thinking about what new product forms might look like — these are things that gradually surfaced in people's awareness this year.

🚥 Koji

We actually discussed this in our last episode of "Crossing" — what OpenAI released during their 12 days of consecutive announcements. Regarding the third point mentioned earlier, about how to get more context, OpenAI also released a new feature: the Mac version of ChatGPT can now read your screen, using on-screen content as context combined with your questions to generate responses.

This screen-reading capability isn't simple screenshotting. It can read content at three levels. The first level is screenshot-style understanding — it comprehends whatever is displayed on screen. The second level is that it can read all content within application windows, even content not currently visible on screen that would require scrolling to see — it can access this information too. The third level is the most impressive: it knows your cursor position. Because where your cursor is often indicates where your attention is most focused. So when you ask questions or discuss things with it, it incorporates your cursor position or selected text into its responses.

So I think this extends beyond just the programming domain — not only the Cursor and Devin mentioned earlier. Even for OpenAI, even within a chatbot format, the application of context makes AI capabilities significantly stronger.

👦🏻 Yusen Dai

Right, the original ChatGPT was somewhat like a pen pal — you could only write to them, one letter at a time. But if this "pen pal" isn't on the other end of email, but standing behind your computer watching how you use it, or even living inside your computer, seeing things that aren't on screen, they obviously become far more useful.

So I think combining AI with user context, with users' existing knowledge, with organizations' existing knowledge — this has enormous impact on AI's usefulness. Because it can now digest so much context, which of course also benefits from advances in model technology.

From ChatGPT to Devin: Four Development Stages and Paradigm Shifts in AI Programming

🚥 Koji

That's not all — Gemini 2.0, released just two weeks ago, also introduced multimodal understanding capabilities. You can simply turn on your camera and point at something, asking "what is this?" I tried it myself: I pointed at a film festival poster on the wall and asked what festival it was, which edition's poster. Questions like these used to exist only in science fiction, but today they've become reality — and this reality operates at acceptable cost, returning answers extremely quickly. Of course it hasn't yet become a polished consumer product, but if people try it out, I think the effect is genuinely stunning.

Let's talk more about AI programming. In the programming domain, there have been very exciting advances this year. Yusen, you've always had strong abilities at framework-building and summarization. Not long ago you shared with me your distilled four-stage theory of AI programming development — want to share it with everyone on the podcast?

👦🏻 Yusen Dai

This actually came from discussions with many friends — it's the crystallization of collective wisdom. AI programming has only been around for just over two years since ChatGPT's emergence, but has already gone through four stages.

The first stage was having AI write code directly, typified by early ChatGPT and Claude. We give it a request like "help me write Snake," and it produces some code. In this process, it neither knows why I want to write Snake, nor how the code runs. I might compile and run it locally, find an error, tell it about the error, and then it gives me a debugged result. At this point, AI was completely like a pen pal you could only communicate with through mail — a simple Q&A mode.

The second stage was represented by GitHub Copilot, where AI began to have context — it could take an entire organization's codebase as context. This gave AI a massive amount of new background information. But users still needed to manually paste code into their IDE for debugging. I consider this the 2.0 stage: we gave AI the codebase as context.

A major advance in 2024 was the emergence of programming copilots represented by Cursor. Its core concept is predicting what code the user will write next. Based on your codebase and what you just wrote, it predicts what code you'll write next, what files you'll create, what operations you'll perform. This involves substantial improvements in both the quality and quantity of generated code, as well as file creation and modification. Later, Windsurf added automation for command-line operations, enabling AI to effectively use my computer. Previously, AI wrote code on a piece of paper and I copied it to run; now AI can create files on my computer, execute command-line operations — entering the "I write for you" stage.

Just when we thought this was already exciting enough, Devin's appearance brought several important breakthroughs: First, it can work asynchronously. While Cursor, Windsurf and similar tools can do quite a lot in one step, they still require sustained attention — "I say one step, it does one step." Devin can work continuously, freeing up the user's attention. This is because it adds a Planner that can plan tasks.

Second, it can execute more operations through virtual machines, doing more debugging work. For example, if you write a website, it can use a virtual machine to visit the site itself, checking whether frontend and backend business logic is correct, and can be interrupted and adjusted at any time. Anyone who's used Cursor or ChatGPT knows you can't make adjustments mid-output — you have to wait for it to finish before modifying anything. But Devin is like a real person: you can give it new instructions while it's completing a task, and it will incorporate this into its existing planner to adjust its plan. This evolves from "writing for you" to "doing for you."

To sum up these four stages: Stage one was AI writing code, represented by ChatGPT. Stage two was AI accessing your codebase, represented by GitHub Copilot. Stage three was AI automatically writing and executing code, represented by Cursor and Windows Terminal. Stage four is AI virtual employees — Devin sets an excellent precedent.

AI Going Global: Deep User Understanding, Clever Content Marketing, Avoid Simple Paid Acquisition

🚥 Koji

Here's a good analogy: AI 1.0 "read ten thousand books" to answer questions, while AI 4.0 "traveled ten thousand miles." It becomes a real employee — you assign it a task, it goes out and completes a full cycle, then comes back to report to you. This is the leapfrog, four-stage evolution we've witnessed firsthand over the past year.

ZhenFund has invested in quite a few AI teams going global, with very typical representatives like HeyGen and Monica, both performing exceptionally well. So I'd like to explore the going-global topic with you.

There's a widely circulated saying in the industry this year: "Go global or go home." Going global seems to have become very important, even critical. So first, I'd like to ask: why is overseas AI adoption so different from domestic adoption? To the point where we encourage domestic entrepreneurs who can barely speak English to bravely give AI globalization a try?

👦🏻 Yusen Dai

I think the core reason is that AI is currently primarily a productivity-enhancing technology, and in regions like Europe and the US where per-capita wages are much higher, people have stronger willingness to pay for tools.

So when you build a productivity tool — like the ones we've invested in, such as HeyGen, Monica, Oculus, Max AI, and others — overseas users, especially in Europe and the US, have a stronger willingness to pay for productivity, and they pay in dollars, so the absolute amount is higher. This is the most important factor.

There are also other reasons: for instance, going overseas gives you access to more capable models like Sonnet 3.5 or GPT-4o, unlocking more application scenarios, whereas models available domestically still have some gaps. Additionally, once a product is built, since large models can handle multilingual input and output, why not promote it globally?

I think the widespread adoption of subscription models is something that's genuinely harder to implement domestically, but has already been widely accepted overseas. This significantly improves startup teams' ability to generate commercial revenue.

🚥 Koji

So what characteristics do you think this generation of AI entrepreneurs needs? And would you encourage them to go global? Because I imagine you wouldn't encourage everyone to do so.

👦🏻 Yusen Dai

Actually, we feel that when all VCs are urging entrepreneurs to go global, that usually means the market is overheated.

We've always been wary of this kind of consensus view. And we believe that for most Chinese entrepreneurs, going global is more of a debuff than a buff — after all, you're playing away games, dealing with problems you don't need to solve domestically and learning about users you never understood before.

So first, I think there are actually many opportunities in China. Companies we've invested in domestically, like Moonshot AI and Yuaiweiwu, are actually growing faster. It's just that monetization may be slightly slower. But I think this is something we learned from the internet era as well. Think back to the internet era — when eBay commercialized early and took commissions, Taobao went free first, and eventually built a stronger business model. So I think China and Western markets naturally suit different business models, and not every team needs to go global.

🚥 Koji

For Chinese entrepreneurs who have already chosen to go global today — and I believe many are listening to this podcast — what advice would you give them?

👦🏻 Yusen Dai

I think going global is just like building products anywhere: you first need to deeply understand users' real needs. In the going-global process, because of language and geographical barriers, this becomes even more important, especially in enterprise services. We've seen quite a few Chinese entrepreneurs in enterprise services who feel their engineering teams are very capable and great at problem-solving, so they believe going global allows them to surpass competitors.

But often, while our teams have strong execution, defining the key problems requires on-the-ground research and truly understanding customers. So especially in sales-driven domains, we believe you must find Go-to-Market experts with relevant experience, and the team actually needs to go to the target destination. For consumer-facing products like Monica, needs may be relatively universal or easier to understand, so that's not necessarily the case. But for enterprise services, I think people absolutely need to go.

Of course, we've seen many doing niche markets (edgy content), because that kind of demand is the easiest to understand — it's probably similar across all humanity. That's the first point: really figure out the users and their needs. Second, I think what successful teams have in common is thinking through and finding a low-cost, high-return marketing strategy. For example, with HeyGen, Monica, Viggle — these Chinese products that have done well going global — they typically excel at SEO, social media distribution, or viral spread of quality content, rather than simply doing paid acquisition.

Of course, if your product has strong monetization, maybe you can make the ROI work on paid acquisition, but basically, paid acquisition is very expensive now.

So how to do marketing cleverly, especially achieving viral marketing through product characteristics, becomes very important.

Using overseas platforms like Twitter well is actually quite different from domestic practice. In China, people are used to buying信息流 ads, doing paid acquisition through sophisticated ad buying. Overseas, I think you need to be more clever about it. Chinese teams generally have strong product execution, so it really comes down to what you build and how you promote it — these two points are where people commonly face challenges, or where doing well can really set you apart.

AI Hardware Entrepreneurship: Looks Beautiful, But Requires Caution

🚥 Koji

There's also a view that quite a few people are building AI hardware this wave, and that AI hardware can particularly leverage China's strengths in manufacturing resources. In the AI hardware space, have you looked at or invested in any projects over the past year?

👦🏻 Yusen Dai

We've looked at quite a few AI hardware projects, but honestly, I think hardware looks beautiful but isn't necessarily that easy to actually land.

What's worked better so far is when the product prototype has already been proven overseas, and we make it faster, cheaper, or smaller. We've also seen teams like Plaud that have created genuinely innovative products. But overall, hardware hasn't scaled that quickly — software remains the more suitable vehicle for AI diffusion right now. So we've always been relatively cautious about hardware.

We have invested in some hardware entrepreneurs, but overall we haven't invested as heavily as some other funds. Personally, I've always been somewhat cautious about AI hardware — including when Rabbit and Humane first came out, I held a relatively skeptical view.

Devin: Not Just a Coding Tool, But the First Genuinely Usable Real AI Agent

🚥 Koji

Alright, let's move to part two of today's discussion — we'll talk with Yusen about Devin. First, a special note: we'll be discussing this from a non-developer's perspective. Neither of us are professional engineers; though we both studied computer science for seven years, we've been product managers since graduation. It wasn't until Cursor launched six months ago that we started coding again — or rather, started commanding AI to write code for us (laughs).

But precisely because of our non-developer backgrounds, we can experience and evaluate Devin from a unique perspective, and predict how AI coding agents — and AI agents more broadly — will transform everyone's future work and life.

Because we believe this generation of AI programming technology will ultimately develop along two directions: one serving professional programmers and developers, the second empowering all non-developers like us. And the latter may have more profound and extensive commercial value and application prospects.

So first question for Yusen: on launch day, you immediately spent $500 to top up. After subscribing to Devin, what was the first thing you used it for? And what's the most impressive thing you've done with it?

👦🏻 Yusen Dai

After installation, Devin suggests some tasks. One of them incorporates your name, searches for information about you online, and builds a personal website for you. After that, I had it do the typical work I usually assign to interns — for example, I gave it a task: we need to revise our venture fund's values statement, called a Manifesto in English. I said, go find out what the Manifestos are for top-tier US VCs. This is a typical task where you roughly know what needs to be found, but it requires information gathering, organization, and problem-solving abilities.

Then I watched it work. There were a lot of interesting things happening. First it had to figure out what counts as a top-tier US VC, so it went to PitchBook, CB Insights, and similar sites looking for ranked lists. It found what it considered the top dozen or so, and when I looked at the list, it was indeed about ten of the most elite firms. Then it went through each of their websites to find their Manifestos.

But "Manifesto" turns out to be called different things at different VCs. Sequoia calls it "Ethos." At Founders Fund it's "Manifesto." Elsewhere it might be "About" or "Philosophy." And a few VC firms didn't have anything like this on their sites at all — no "who we are, what we believe" statement. So I watched Devin trying to understand the task, trying to find the most relevant content.

For example, when it looked at Accel (also a very well-known US VC), it found nothing on the website. So it went into their News section, searched around, and found an article from two or three years ago where they discussed Accel's values and methodology. It pulled that content as what it was looking for. So you can see it solving problems like a junior human employee — not mechanically checking whether your website has something literally called "Manifesto" and giving up if not. Instead, it said, I need to look across your entire site for content that fits this description, and search for that.

It eventually gave me a Markdown file with Manifestos for ten VCs, but it had many of the typical issues you see with current AI models. For instance, it sometimes takes shortcuts. I asked it to pull the full text, but for a few of the VCs, it just summarized the content for itself. This is something we often encounter with AI chatbots now — because of token limits, it won't give you the full text, just an abbreviated version. At that point you have to tell it: give me the complete text. So it really does need teaching, just like an actual intern. But I found its planning ability, and its creative problem-solving when a task couldn't be directly completed, very impressive.

Of course this probably isn't the typical use case for Devin, since I wasn't having it code — I had it do something that language model AIs commonly do. So I can easily imagine: now that we have Devin for programming, we could absolutely have corresponding Agent products suited for text-based work, for finance, for legal work.

I believe that as long as the work I've defined is something a person can do sitting at a computer — using a computer, going online, using software — then it can more or less be performed in this workflow.

That really impressed me.

🚥 Koji

So what I want to ask is: from day one until now, and it's not been long, about two weeks, what kind of future do you feel you've experienced?

👦🏻 Yusen Dai

After using Devin, I feel that as the first truly usable Agent product, it may mark an important moment in human history.

Why do I say this? Because humans have invented many tools, and some say "humans are the animal that can use tools." But essentially all these tools fall into two categories. The first are tools requiring sustained attention — like a power drill, a hammer, or a keyboard and mouse. They need our continuous focus and input. The second are mechanically repetitive automated tools — like washing machines, vending machines, assembly lines. They don't need our attention, but can only handle repetitive tasks.

We've always been searching for a third kind — something that doesn't need sustained attention, but can plan and solve problems on its own. This is what's called an Autonomous Agent.

In the original vision, perhaps only hardware products like Viggle achieved this. At the software level, we had never seen such a product emerge. Last year there were attempts like AutoGPT, but they all remained at the prototype stage.

I found that Cursor defined several characteristics that a true Agent product needs:

First is the asynchronous experience enabled by powerful task-planning ability. Devin's originally designed scenario was: in Slack, you @ Devin saying "help me fix this bug," and it just goes and fixes it. It only comes to me when it genuinely needs help or has completed the task. This is very much like an intern — you assign the task, they work on their own, and only come to you when they encounter something they can't solve. Meanwhile, I can assign tasks to multiple interns, letting me focus on what truly matters.

Second, it runs on a cloud-deployed virtual machine, so it can use a browser, and in the future will be able to use more software, thereby completing more tasks. This is completely different from Cursor and Windsurf, which run on my own computer. If you've used RPA software before, you'll find that when RPA is operating, you dare not touch anything, because your actions will interrupt its workflow. The AI is using your computer. But Devin uses a virtual machine — just like how we provide interns with computers. The flexibility that AI gets from having its own virtual machine is fundamentally different.

Third, Devin learns and grows like a real employee while working. For example, when we hire an intern, they'll definitely mess up a lot of things on day one, because they don't know how to handle many social behaviors within our organization. As they do things, they'll gradually realize they need to accumulate relevant knowledge. In this process, these experiences are called knowledge. It proactively prompts that it's learned some knowledge point — for instance, when searching for information, try to look on official websites first. I'll confirm that it's learned these good practices. This process is very similar to doing reviews with interns or employees. Just like when an employee writes a work summary saying what they've learned, we affirm "yes, those points are correct." In theory, this allows continuous accumulation of organization-specific proprietary knowledge, making it increasingly adapted to this team.

Actually, this is how we hire too. When an employee first arrives, their value is relatively limited; they need continuous learning to better adapt to the organization. But previously when using tools, we expected the tool to work perfectly the moment we opened it — we wouldn't expect a computer to keep learning before it became more useful.

With Devin, we're truly seeing something with a growth curve similar to human employees. Though still early, we find this paradigm shift very important.

Fourth, Devin introduced a pricing model based on task completion. $500 corresponds to 250 ACUs, with each ACU being about 15 minutes of work, which translates to roughly $8 per hour. This is already less than half of California's minimum wage ($16/hour). As AI compute improves and costs decline, this investment will accomplish even more in the future. Compared to hiring people and dealing with HR, office space, management, and other issues, AI is a 7×24 tireless employee.

A friend put it quite well:

Programmers like Cursor because it's a programmer's Copilot, helping improve efficiency; bosses like Devin because bosses think about how to spend money to buy productivity. Devin shows a potential paradigm shift — expanding productivity through spending money. I think Devin showed me the Scaling Law of work.

In many Coding Agents, the first task is often building a personal website, and we joke that "this is the new Hello World." Devin did well on this task because finding my information online is relatively easy, and it could quickly build the site.

🚥 Koji

So Devin's emergence doesn't just make people feel that AI programming has become impressive — it defines a new interaction paradigm. Everyone can see how an AI Agent can work this way. Because Yusen and I use a team account in Devin, I can see all his task progress, see how he uses Devin, and how Devin responds to him.

Like the task just mentioned, one supplement is that immediately after receiving instructions, Devin will first tell you its work plan. It manages upward like reporting to a boss, saying first I need to understand this task, I need to break it down, I'll divide it into several parts to execute, and then it proactively reports back after completing each step. When it encounters situations where it can't proceed, it will also tell you and ask for your guidance — this is quite stunning.

The second point is that Yusen's task had a follow-up. After pulling the Manifestos from ten top VCs, Yusen had it build a website. It spent an hour on the first version, which was quite rough. Right at that time I logged into the team account and saw the report it had submitted. I thought, okay, let me take over and assign this task. I gave it some new instructions — for example, I gave it a reference website, said this style looks good, have it adjust the website following this style. At the same time I wanted to try Recraft's illustration generation API, so I threw the API documentation and key to it, asking it to create an illustration for each of these ten Manifestos.

What I want to express is, at that moment there was a feeling of really being in an office. There was an intern initially helping Yusen with work, but now it had produced a report, and right when Yusen had gone downstairs to eat. Then I saw its report and gave it some advice: actually, what Yusen wants is this, go refine it further, and when he comes back he can review it. This feeling of really using a person is why we say it's a true Agent. Because "Agent" translates to "person" — not just machine, it carries the meaning of some kind of assistant. This is why I feel Devin has created a new paradigm, like using an assistant.

👦🏻 Yusen Dai

Right, there are many interesting details here. Let me give another example. In another friend's task, he had Devin scrape some people's information from LinkedIn. For instance, OpenAI's China-based employees. But Devin obviously didn't have a LinkedIn account, so it needed to ask the user: could you help me log into LinkedIn? At this point, because Devin runs on a virtual machine, it has an interactive mode. As the user, I can enter my LinkedIn username and password into the virtual machine, and then Devin continues working.

What is this like? For example, we hire an intern, provide them with a computer, but they don't have a subscription account for specific software, so they'll say "Boss, could you enter your account?" After I input my account, they continue working with my logged-in account.

That's why the virtual machine becomes so important — it can handle all these operations inside without interrupting my workflow.

Otherwise it's like Cursor or Windows borrowing my computer, where I can't get anything done myself. This asynchronous approach lets me assign Devin many tasks at once. It's a parallel working mode — I just need to pay for compute costs.

This actually matters a lot. In daily life, having one intern is one thing. But if I had ten interns, each capable of helping with many tasks, the productivity gains could be exponential.

🚥 Koji

Using Devin reminds me of when people used to say "everyone is a product manager." Today it's become "everyone is a CEO." Because in interacting with AI agents, you seem to only do the three things CEOs love most: first, give orders; second, check the work; third, at a more sophisticated level, offer inspiration and guidance.

👦🏻 Yusen Dai

In fact, many people using Devin or other AI products encounter the same problem: what should I do, and how should I articulate my needs? Imagine if we hired an employee and simply told them, "Build me a Taobao" — that person definitely couldn't deliver. So why do we often have unrealistic expectations of AI, thinking we can say "build me a Taobao" and it'll just happen? That's clearly wrong.

Indeed, each of us needs to think about what we actually want to do. When facing a powerful model with abundant intelligence and capabilities, the key question is whether you clearly know what you want to achieve, and whether you can express that need in a more reasonable, comprehensible, and structured way.

Just as we ourselves get frustrated as product managers, designers, or programmers when dealing with bosses who haven't figured out their own requirements — like asking for "colorful black." But when we become the boss of AI, can we be a good boss? This is what everyone needs to learn going forward: how to be a good boss.

🚥 Koji

There's another strong impression from using it, something hidecloud pointed out recently. He reminded everyone that one of Devin's most powerful aspects is its ability to help us tap into the accumulated wisdom of human history. What does this mean?

When we need to accomplish a task, we often don't know that a wheel already exists, that someone has already built this tool. Because many tools exist as code, as repositories on GitHub or Hugging Face. Downloading such code locally, deploying it on a machine, and connecting it to run with other work or software — perhaps only one in a thousand people could do this. But today with Devin, theoretically anyone can, because you can give instructions in natural language like a boss.

A concrete example: say we want to build a chess application. In the past, just writing out the rules of chess would take hundreds or even thousands of lines of code. You might think to search whether someone has already written these rules as a callable code library. But you might get hundreds of pages of Google results, with no way to know what's best or what represents best practices. With Devin, you can give it this command, and it will use its own analytical approach to find the most suitable existing code library and put it to use directly.

The value this brings: all the tools or code libraries that previous developers built to solve specific problems become directly usable — no need to reinvent the wheel. You can stand on the shoulders of giants, using these community-validated best practices to build the tools you want. I think this is a value that Devin and Cursor are realizing — perhaps not the most visible, but profoundly impactful.

👦🏻 Yusen Dai

When ChatGPT first appeared, I had a strong intuition: if much of your work involves copy-pasting or being a "frankenstein" stitch-together job, that's easily replaceable. People found that the earliest work dramatically made more efficient by AI — or to put it less kindly, most easily replaced — was entry-level design work of the cut-and-paste variety. Like copying someone else's design, or entry-level coders taking some library, making simple modifications, and applying it to their project. This kind of work is most easily replaced, so frontend programmers actually face significant pressure, because frontend presentation mostly doesn't require that much innovative thinking.

In this process, I think for everyone, the abilities to generate ideas and solve problems creatively will become increasingly important.

Finding existing solutions and gluing them together is precisely what AI excels at. Most of what we do at work involves problems that have already been solved, or wheels that have already been invented — it's just that humans previously didn't know these wheels existed, or couldn't stitch them together well. But now AI can help us do this, allowing us to focus on thinking about "what to do" — and this will become increasingly important.

This also makes me think about the impact on education. So much of our previous education and training taught "how to execute." It's like when there were no calculators, we had to learn extensive manual and mental calculation. But now, we need to understand the principles of computation without necessarily performing the calculations ourselves. We can devote more energy to thinking about what to do and asking the right questions. This is why I believe future education systems need fundamental transformation.

🚥 Koji

So 2025 is something to look forward to. From Devin's release, what we see isn't just AI Coding being elevated to the next level by Agent technology — across the board, whether in law, business analysis, or education, the emergence of this new paradigm will bring disruptive revolution, and with it, entrepreneurial opportunities.

Earlier, Yusen mentioned a fascinating point: Devin is the first tool in human history that requires neither sustained attention nor mere mechanical repetition. This also reveals something like a Scaling Law of work. Could you expand on this? Help everyone better understand what remarkable value this represents.

👦🏻 Yusen Dai

First, on Scaling Law — the most straightforward explanation is that I can get more productivity by spending more money, where money equates to compute. This is actually quite remarkable. Think about it: many companies raise lots of money but can't effectively convert it into productivity — they need to hire people, build organizations, handle all sorts of trivial matters. But with these asynchronously working AI agents, we can assign many tasks to different types of AI to execute. They consume compute and electricity to complete the tasks themselves, and can work in parallel.

You can easily imagine a "product manager-type" AI that's better at articulating and breaking down requirements, directing many AI programmers to work, forming a virtual organization.

In this organization, you need to focus on two things: first, what do you want to do; second, having sufficient compute and capital investment. In this kind of organization that's rapidly becoming reality, we can effectively scale work through more money and compute. This is what's called the Scaling Law of work.

The second point is interesting. We often hear entrepreneurs say, "I have a great idea, but I'm missing a programmer."

Excellent programming execution capability is still a scarce resource. But when execution itself is no longer scarce, "what to do" becomes paramount.

As mentioned earlier, everyone needs to learn to be a boss. This way we can see more entrepreneurial opportunities — many entrepreneurs whose ideas were previously buried due to lack of excellent programmers may now get more chances, more creativity may be put into practice. This is also why we can scale entrepreneurship itself: because productivity can be increased through capital investment.

All this is possible because AI agents can work in parallel. If our attention had to be on the tool itself, attention is limited. But now our attention can be distributed across different agents — one person can simultaneously give instructions to multiple agents to complete tasks.

🚥 Koji

Speaking of Scaling Law, I'm reminded of an analogy. Years ago, Xing Wang had us read a book called The Leadership Pipeline. The book discusses an important cognitive shift when you first become a leader of a small team: your output is no longer your personal output, but the output of the entire team.

Today, the Scaling Law of work we see in Devin is similar. The output here is no longer what you personally produce while focused on the task at hand — it depends on how well you delegate team tasks and set inspection standards. All the team's output, including all of Devin's output, ultimately becomes your output. This means you can achieve unlimited scale-up with limited attention. As long as you can manage enough people and agents, and managing AI agents is much easier than managing people, because managing people involves more communication, coordination, and emotional labor. I think this is what Yusen meant by the Scaling Law of work.

👦🏻 Yusen Dai

Right, Xing Wang recommended the book The Leadership Pipeline. The concept holds: imagine if you could become the CEO of a multinational corporation, commanding thousands or tens of thousands of people — what could you accomplish? We didn't have such opportunities before, but now we can gain similar capabilities by managing AI agents and having agents direct other agents. What this requires is money and compute, and many companies aren't actually short of money — they're short of talent, of organizational structures that can execute.

So I believe two trends will emerge: on one hand, capable companies and individuals will be able to do more; on the other hand, many people with ideas will be able to quickly realize them at relatively low cost, gaining user validation or investment — giving us more entrepreneurs and more room for innovation.

🚥 Koji

Right, this is one of the most popular phrases this year: "super individual." Because as individuals gain empowerment from more and more tools, including AI agents, they can accomplish what previously required ten or twenty people.

However, shortly after Devin's release, it received considerable criticism and complaints. What's your take on that?

👦🏻 Yusen Dai

Much of the criticism focused on the $500 price point, comparing it to Cursor's $20 price. First, I believe these are two fundamentally different paradigms.

One is a tool that requires my time to use. It makes my time more efficient, but it doesn't save time. So when using a tool-type product like Cursor, my costs don't actually decrease — it's my costs plus the tool's costs. But if you treat it as an employee, the comparison becomes an employee's salary. As long as it can do more work than an employee you could hire for the same price, I think this price point is acceptable in the US and European markets. Many people see the price and immediately ask if it's a rip-off, but the key is how you view and use it.

I've discussed with programmers their experiences using Cursor and Devin, and found that when Devin's capabilities were still limited, using it represented a major shift in most programmers' workflows. Because programmers understand how code runs themselves, they often want to maintain control over the big picture. So at this stage, a Copilot like Cursor is a better fit for their current workflow. Programmers already accustomed to working in an IDE, when they have tasks to complete, find that having to talk to Devin, wait for Devin to work, and then review the results isn't very efficient. They'd rather fix bugs or write code themselves. So if you're a very skilled programmer, you probably wouldn't want to be stuck with a capability-limited intern — because Devin right now is only at an intern level, and training an intern takes time and patience.

At this point programmers might think, rather than waiting for you to write code and then helping you solve problems, I might as well write it myself. I think this is completely understandable at an early stage of technology — we need to look at this from a human perspective. When a person makes mistakes, as managers we tend to be more patient, because we know people can learn and grow. Point out their problem today, they might remember it, then have more motivation to work, and through training become a decent programmer.

But Devin can actually learn. Yet we haven't yet established the expectation for AI software and products that "it can grow, it can learn, it can be managed."

So when it has problems, many users' reaction becomes "I bought this $500 tool, and it still has issues?" — feeling disappointed. So when introducing a product like Devin to an enterprise, managing expectations becomes very important. Including Devin itself notes in its documentation that it first does tasks you would assign to an intern — simple frontend tasks, fixing bugs, adding a Dark Mode toggle to the frontend, that kind of work.

But the ability to ask good questions is also something humans need to learn. I often see people make requests like "help me build a Taobao" or "help me make a WeChat" — this far exceeds its capabilities. Right now Devin, like all AI products, will dumbly accept this task and say "okay, I'll help you build a Taobao." The results in this case definitely won't be satisfactory. So I think learning to use a tool well requires learning — we're not yet at the stage where any request can be directly completed. That wouldn't be an intern, that would be a god.

As Devin's capabilities improve and its understanding of organizational environments deepens, I believe it will gradually grow from intern to junior full-time employee, then to senior full-time employee. This requires a process of acceptance.

I think Cursor is incremental innovation on existing workflows — it doesn't fundamentally transform programmers' work. But Devin represents a disruptive innovation logic, which often requires much adaptation time and a different onboarding process. The first product may not necessarily achieve this, so I don't think Devin is necessarily the final answer.

It's very possible Devin is just showing one form that future AI products might take. Truly learning to adapt to and use AI-type products, like adapting to the concept of SaaS, or adapting to distributed work concepts like remote work, all require long periods of time and suitable catalysts. So I think it gives us great directional guidance, but it's still at an intern level right now. It's easy to point out its problems in this process, but more important is that it proposes this future direction — getting inspired from here to build better agents, I think that's the key.

🚥 Koji

This is like the half-full glass theory — some people see value in the half glass, others see problems. Just as we discussed earlier when Devin completed the task of "finding the manifestos of ten top VCs," it knew how to find this content from press releases when the Accel website lacked relevant background. This is a huge highlight — it sets tasks, reflects, self-checks. On the other hand there are indeed many problems, like the webpages it produces being very unattractive. But seeing highlights rather than problems, seeing future possibilities rather than points worth criticizing right now — this makes me think: critics often feel correct, but only builders, though they may appear clumsy, are more likely to succeed.

This reminds me of something Huiwen Wang once said: If you believe something will eventually happen, do it once every three years. Agents have been believed to appear ever since humans had science fiction, and people occasionally try. But after seeing Devin, it feels like this may be the closest we've come to success.

Let's talk about 2025. Throughout 2024, although our discussions were fairly optimistic, the broader environment would occasionally produce various pessimistic narratives. I especially remember in the second and third quarters, the entire discourse was discussing where AI's PMF actually was — it seemed this wave of AI落地 was harder than expected.

Now standing at the beginning of 2025, there's a very simple yes or no question: Yusen, are you optimistic about 2025?

👦🏻 Yusen Dai

I'm actually still quite optimistic. First, I don't think we should expect AI applications to find PMF that quickly.

I often use an analogy — although many people compare ChatGPT's release to the iPhone's release, saying AI has entered the iPhone era, I consistently believe it represents more of a BlackBerry era.

What's the difference between the BlackBerry era and the iPhone era? Many listeners may not have used a BlackBerry — this belongs to us post-80s memories. Before the iPhone launched, smartphone form factors were very non-unified, because the technology was still relatively early, development was fragmented, and people hadn't found a convergent path. This meant many things were wanted but couldn't be done, the technology itself was expensive, there were no unified development standards or product standards, and there were few developers. So at that time, wanting to build truly popular applications on mobile internet, like Douyin, was very difficult to achieve. I've repeatedly made this point: you couldn't build Douyin in the BlackBerry era. As technology advances, moving from the BlackBerry era to the iPhone era unlocks more application opportunities.

After the iPhone appeared, first the technology became good enough that many applications went from "wanted to do" to "could do" — including good cameras, good screens, good processors. Second, technology became standardized. After the iPhone launched, phones all looked the same, and people realized the technology direction had converged. At the same time more developers emerged, because development became easier, technology became standardized and cheaper, and people understood it better — so the iPhone era gave birth to massive numbers of applications.

When ChatGPT first came out, we also found many things we could imagine but couldn't achieve. Agent is a typical example — in the first half of 2023 there was an attempt called Auto-GPT, which proposed many good concepts, also using language models to first make plans, then check completion and iterate. But models at that time had too many hallucinations, couldn't effectively use tools, and couldn't effectively browse the web, so it simply couldn't work. This is a typical example of "trying to build Douyin in the BlackBerry era."

Now as Agents have progressed in reasoning ability, programming ability, and tool-use ability, Agents look much more the part. Although there are still many shortcomings, at least they've reached the first step of being usable at an intern level. This is a typical example of technological progress unlocking more application opportunities — I believe this is one that will eventually take us from the BlackBerry era to the iPhone era.

From ChatGPT's emergence to now, these two years, we've seen enormous progress, which makes me very optimistic.

In just two years, AI programming has gone from ChatGPT's "you ask, I answer" to Devin's "you ask, I do" and Cursor's "you ask, I write" — bringing very significant progress, and this speed is actually quite fast.

Second, I think PMF often comes from technological progress itself. For example, the product Cursor actually appeared in 2023, but its proposed prediction of next actions required more powerful models to predict and write better code. You could say Sonnet 3.5's emergence enabled Cursor to truly deliver the product experience it wanted. Sonnet 3.5 activated the product experience Cursor wanted to deliver, and Cursor's popularization also quickly made Sonnet 3.5 the most popular model in the AI programming space — this is a mutually reinforcing relationship.

Similarly, for a product like Devin to succeed, it also needs models to improve in reasoning, tool-use, and other capabilities. Sonnet 3.5 or GPT-4o right now may not be sufficient to do it well. So the Devin product form may need a more advanced model to activate it — this model could be o1, o3, or another new Anthropic model. This is a reciprocal process where a product waits to be activated by a model, then enables the model to be widely used — so this stage indeed requires technological and model progress.

The mature mobile internet period we just experienced had a characteristic where products were very easy to use. For example, Douyin just takes moving your finger, WeChat and Xiaohongshu are all very easy to pick up. But when we come to an early stage of technology, using a product well requires some threshold and learning. Think about the earliest smartphones, personal computers, the internet — they all required learning to use.

Right now when people use AI, they're far from extracting all the intelligence in the product. Current large models, whether OpenAI, Claude, or Moonshot AI, have actually compressed massive amounts of knowledge and intelligence. But have we learned to correctly use it, to efficiently prompt, to efficiently extract the intelligence in the model?

I think most people haven't learned, including myself. I'm constantly discovering that the model can actually do this for me, answer this kind of question. So in this process, we're transitioning from the easy-to-use-product mobile internet era to the deep AI era where learning is required to use products.

At this time people's initial experience will be somewhat frustrating, feeling the product is a bit difficult to use — this is the characteristic of an early technology stage.

A lot of the time, the applications can already do many things — we just haven't learned how to use them yet. We haven't become good prompters or good managers.

All of this requires learning, or requires waiting for models to become powerful enough to help us do these things. By then we may enter another period of product application, but right now products are still in a stage of磨合 [working out the kinks] with us.

🚥 Koji

So everyone needs to understand where the boundaries are through trial and error, and how those boundaries keep expanding. I want to add something — beyond the new opportunities unlocked by technological and model advances that we just discussed, especially in the agent space, there's a fourth dimension.

In the previous episode of Crossing, when we discussed OpenAI's 12 Days of OpenAI, our guest Dacongming mentioned that there was some heavyweight content the company didn't announce, whether out of PR considerations or to avoid drawing excessive attention from competitors. One point that's crucial for agents is OpenAI's current function calling and structured output capabilities, which allow agents to receive much more precise instructions. This was probably overlooked before, but once you say it out loud, it makes a lot of sense.

Looking ahead to 2025, Yusen, what application directions do you think are relatively easy to land? This is probably something entrepreneurs are paying close attention to right now.

👦🏻 Yusen Dai

Looking at what's been easier to落地 [commercialize] over the past two years, I think there are a few. First: anything that helps customers make money. Of course, if your technology isn't fully mature yet but can directly help me make money, or directly improve efficiency in the commercialization process, that becomes very important. Take Midjourney, for example — it has hundreds of millions in annualized revenue, with roughly half coming from advertising-related demand, meaning people use it to generate commercial images for ad campaigns. This is a very practical scenario. I was already making these ads to make money; now I can produce ad content faster and better. Or Heygen, which is also primarily used in marketing scenarios — people use it to create promotional video content. So first, technology that helps customers make money — in the early stages, people are willing to spend time learning and figuring it out.

Second: anything that can improve productivity by tenfold or more on important tasks. Because if a good technology only improves productivity by 50%, people will still face a lot of resistance. It has to bring extremely strong productivity gains — something like Cursor or Devin, which are absolutely tenfold productivity improvements for programmers. Programmers might spend ages searching through codebases, so their motivation to use these tools becomes very strong.

Another example is AI search engines like Perplexity, which I think also represent a tenfold productivity improvement over traditional search. Because before, if I wanted to find information about Koji, I'd have to search through lots of content and read a dozen or two dozen articles from The Fair. Now I just ask it, and it reads those dozens of web pages for me and summarizes them. So for information gathering and question-answering, it's more than ten times more efficient than search engines. I think products like this can relatively easily find product-market fit.

Third: satisfying basic human needs, like NSFW content — we've all seen plenty of scenarios like this. Overall, either it can make money, or it can help me achieve very high efficiency. Achieving either one is excellent.

🚥 Koji

What application directions do you think people should be somewhat cautious about, that are a bit difficult to pull off?

👦🏻 Yusen Dai

In mobile internet, many of the winners were "time-killing" applications. In China, people got used to building apps with high user stickiness, where users spend lots of time, and then monetizing through advertising. ByteDance, Xiaohongshu, and Kuaishou all follow this paradigm. This was the existing paradigm of mobile internet because it was a new device that made previously unconnected time usable — a zero-to-one logic.

Now that apps like Douyin already occupy so much of our time, if an AI application tries from the start to compete with these mature players on "time-killing," it runs into competitors that are already extremely powerful and have already captured most of people's time. Doing "time-killing" apps at this point is very difficult.

So what we've found is that the only things that can actually work out are relatively niche, audience-specific products. AI companion chat for general users is hard-pressed to be more attractive than video apps like Douyin. So be cautious about applications that compete with giants for users' time.

Second, changing the physical world is still quite hard. We just talked about AI writing code, AI using tools — these are all in the digital world. AI can do many things in the digital world, but in the physical world, even basic actions like picking up a cup are still quite difficult for AI.

Although humanoid robots are extremely hot right now, in this direction, the technical implementation path and how to scale model data — these are still open questions without clear answers.

So I think for the next three to five years, applications that change the physical world will still face many challenges.

Third, over the past couple of years, quite a few devices have tried to replace the phone, like Rabbit, Humane, and others. They emphasized building a phone replacement product, and now there are roughly 100 teams working on smart glasses.

My view is: if the scenario you're doing is one that the phone already covers — like making calls, searching for nearby information, listening to music — then replacing the phone is extremely difficult.

So far, hardware that can coexist with phones is basically doing things phones completely cannot do. Drones can fly, smartwatches can be worn on the wrist, smart rings can go on your finger, or like Insta360 which can be used in sports scenarios.

But products like Humane and Rabbit are actually doing scenarios that phones already do well. At that point, users have very little motivation to switch, because phones can already achieve at least 80% in most scenarios. Unless what you're building is vastly, vastly better, or something phones fundamentally cannot do, replacing the phone will be very difficult.

For Agent-type products, I think in 2025 we'll see an enormous number of Agent products emerge. Many of these will face a challenge: when you need to make major changes to an organization, can you actually achieve that? For example, Devin faces the challenge of changing how programmers work, from writing code themselves to directing others to write code. This kind of workflow change faces a lot of resistance in many organizations, especially large companies.

What we're finding now is that pushing AI in large companies also involves many issues around data permissions, privacy, and security. If you're changing workflows, many people's jobs will change, which creates even greater difficulty. So I think making major organizational changes — unless you can significantly improve productivity to the point where the organization has no choice but to adopt it, or unless you're targeting small and medium-sized businesses.

Otherwise, making major changes in large organizations often faces barriers of human nature, not technical barriers.

2025 Outlook: Agents, Personalized Services, and Superhuman-Level Breakthroughs

🚥 Koji

We just discussed how technological unlocks bring new opportunities. We've talked quite a bit about model reasoning capabilities, reduced hallucination, and tool use like Computer Use bringing agent opportunities.

Beyond this, what other technological unlocks do you think could bring wave-like AI entrepreneurship opportunities in 2025?

👦🏻 Yusen Dai

I've summarized a few technological unlock directions that could bring wave-like AI entrepreneurship opportunities in 2025:

The first is Agent. We discussed this earlier — we'll see AI products emerging for various domains. They'll draw on Devin's approach, doing asynchronous tool use and charging by workload.

In the United States, some people have flipped the original SaaS (Software as a Service) to "Service as Software" — turning services into software sales, or "sell work, not software," selling work outcomes rather than the tools themselves.

2025 may see many such attempts. Many will fail, but some interesting products will emerge.

The second is "Scalable Personalization." Looking back at the evolution of internet content distribution: first came portals with "one face for a thousand people," where everyone saw the same thing; then search engines, which personalized content to keywords but gave the same results for the same keywords; then recommendation algorithms represented by Douyin, which proactively pushed content based on user context.

Now we're thinking about the next step in personalization: if the content a user wants to see doesn't yet exist, generate it for them. For example, video generation technology like Sora is about generating content according to personalized needs. Recently fast-growing applications like bolt.new and Windsurf both generate personalized websites through text prompts. In software development, the future may no longer be "Hollywood blockbuster-style" centralized development like WeChat or Douyin, but rather more personalized software/content experiences for each user category.

Google's NotebookLM also reflects this trend. For example, with podcast content, right now we can only listen to already-recorded conversations, but in the future AI might generate conversations between any two people on specific topics. As AI capabilities improve, the software we use and content we consume will all become more personalized.

The third is that in o3, we can see AI capabilities evolving from "ordinary human level" to "superhuman level." Early tests like MMLU were still evaluating whether AI reached ordinary human levels; now they've shifted to benchmarks targeting elite humans, such as SWE-bench for programmers, the American Invitational Mathematics Examination (AIME) for high school math competitions, and the Graduate-Level Google-Proof Q&A (GPQA) for PhD qualifying exams. In early 2024, advanced models like o1 and o3 had already reached roughly 80% on these tests.

We now need to establish superhuman-level benchmarks, such as FrontierMath endorsed by Terence Tao. o3 recently scored 2700 on Codeforces, a level that only about 130 people in all of humanity have ever reached. This means AI will play important roles in scientific research and frontier exploration.

So when I saw o3 come out, some people criticized it for being expensive to run a task, with high compute costs.

But what I want to say is, o3's high-compute mode was never meant for ordinary tasks. Its purpose is to tackle the hardest frontier research and exploration problems facing humanity. It's entirely normal for something like that to be expensive.

In fact, we'll likely find that AI models diverge between everyday tasks and frontier research. It's like Sheldon in The Big Bang Theory — a brilliant scientist who's completely hopeless at daily tasks.

Some AI models will be more like Sheldon, solving frontier exploration problems; others will be like the affordable, capable o3 mini, mainly for getting work done, more like a programmer; and there will be even simpler models just for answering basic on-device questions, like "What's the weather today?"

So in this landscape, we can see everyday needs being solved ever more efficiently and cheaply, while at the same time, in genuine frontier research, AI collaborates with scientists to produce new advances and generate new knowledge for humanity. That excites me tremendously.

🚥 Koji

Because this year also brought significant breakthroughs in multimodality — whether it's 4o realtime voice, or what OpenAI released in a rather inconspicuous corner of their event but which many considered one of the most noteworthy achievements of those 12 days: their many-to-many multimodal interaction.

What entrepreneurial opportunities do you think multimodality will bring next year?

👦🏻 Yusen Dai

In multimodality, I think the first important thing is how AI understands this multimodal world.

For text, take a simple sentence like "The weather is nice today." It's a very straightforward string of characters, but it contains a great deal that requires seeing to truly understand — so a picture is worth a thousand words.

Images and video contain enormous amounts of information. If AI cannot adequately comprehend this information, its intelligence will have major gaps. Current AI is like a blind person — a blind person can still solve incredibly difficult math problems, which may not hinder much, but to possess more complete intelligence, multimodal understanding capability is genuinely important.

OpenAI and leading overseas researchers generally believe generative capability may not be the most important thing, which is why Sora now receives relatively few resources. In the United States, multimodal generation is a relatively parallel track because its application scenarios are mainly entertainment content and content production, so it still seems somewhat distant from AGI. Companies like Anthropic, which don't do multimodal generation, believe AGI can be achieved through text, code, and APIs alone — a different perspective.

On the topic of multimodality, I think NotebookLM gives us excellent insight: how to transform content from one modality to another for consumption.

For example, the TTS we used to do converted text directly to speech, but turning text into a podcast isn't simply reading it aloud — that would just be audiobooks. Podcasts require transforming content into a form better suited for audio consumption. Similarly, text to video: when we adapt Romance of the Three Kingdoms into a TV drama, it's not simple reproduction but requires artistic adaptation. Video to text, video to audio — it's the same principle. Naturally converting between different modalities and creating content best suited for consumption within each modality is an incredibly exciting process.

Suppose I enjoy scrolling Douyin — I could transform The Three-Body Problem into content suited for Douyin consumption, or into content suited for podcast consumption. This opens up many opportunities in content consumption.

Going further, people believe multimodal generation and understanding will be tremendously helpful for embodied intelligence. We're seeing much frontier research, like the recent Genesis project, studying how to simulate the physical world and how robots manipulate real-world objects — all fascinating research. Though I haven't been following this area as closely recently.

Overall, conversion between modalities is indeed a very important direction. As you mentioned with Gemini 2.0, it can efficiently understand received video signals. This enables some very intuitive application scenarios. In daily life, there are many things we see but don't know how to use — but if its video generation capability is strong enough, it could directly overlay usage instructions on the video feed. For instance, we previously discussed a scenario with Google researchers: I have a coffee machine at home, I point my phone at it, and the video stream directly overlays a generated video prompt saying "Press this button to start brewing coffee." The instructional video is generated but superimposed on the existing video. These are all fascinating ideas, though they may require further technological advancement.

AI Native Applications: Waiting for New Business Models After Deep Technology Diffusion

🚥 Koji

I think 2025 will likely see such applications emerge, including their combination with AI hardware. For example, I previously saw a demo of playing tennis with AI glasses — it gives you real-time coaching, telling you how to adjust your posture and receive the ball when it comes from your opponent, helping you improve your game.

On many-to-many interaction, I want to expand a bit — this is a development that's genuinely surprised me recently. As a guest mentioned on the previous "Crossing" episode, this technology was released during the 12-day event but placed in an inconspicuous corner. He considered it actually the most noteworthy breakthrough. OpenAI chose to reveal this information quietly to avoid drawing competitors' attention. However, within the developer community, they still conducted one-on-one outreach with key developers.

What's special about this technology is that it can simultaneously receive multimodal input and simultaneously output multimodal content. And this input and output is many-to-many. Everyone understands the end-to-end concept — many-to-many is essentially several orders of magnitude leap beyond end-to-end.

I also want to ask Yusen an interesting question that everyone must be curious about: What do you think the big opportunity for AI native applications might look like?

👦🏻 Yusen Dai

First, I believe major opportunities will emerge after deep AI technology diffusion. If it's still only being used by niche populations, the big opportunity may not yet be visible. Let's review how internet native applications and mobile internet native applications historically emerged.

Step one: as technology diffuses, use new technology to solve old problems. For example, in the internet era, we had email solving communication, portals solving news consumption, and self-operated e-commerce solving retail. But as the internet expanded further, social networks only appeared once everyone was online; search engines only became necessary once information was online; and platform e-commerce only emerged once buyers, sellers, payments, and logistics were all well-established. These platform e-commerce businesses, social networks, and search engines were the true internet native applications — all created by startups, and ultimately occupying the largest market capitalizations.

Mobile internet native applications followed similar logic. Once mobile internet (including smartphone hardware and 4G networks) became ubiquitous, content producers and consumers all had smartphones — only then could mobile information platforms like Douyin, Kuaishou, and Xiaohongshu emerge. Once blue-collar workers all had smartphones, applications like Meituan delivery and DiDi could be born; once gamers all had phones, mobile native games like miHoYo's titles and Honor of Kings could appear.

AI native applications should follow similar logic. Initially there may be applications like ChatGPT, giving everyone an AI assistant, but its diffusion scale needs to grow larger. Once each of us has our own AI assistant, using AI to solve many work problems, even conducting meetings as we are now, new possibilities will emerge.

What happens when AI interacts with AI? For instance, in a company where most work execution is completed by AI, this could produce enormous changes in productivity and enterprise service software. Because you not only need to execute, but also manage these AIs, assigning them tasks and breaking down those tasks. These may be things humans simply cannot do, because humans don't have that much attention and energy.

Another important theme is monetization in the AI era. In the mobile internet and internet eras, massive monetization happened through advertising. But when you ask Kimi or Perplexity a question, the ads in search engines and on webpages aren't seen — because AI is reading those webpages for you. This requires reconstructing value capture mechanisms. How do we extract value from the answers I get from AI? Originally ads were for humans to see, but when AI sees ads, it filters them out. So the disruption of advertising business models will also create many opportunities for AI native applications.

🚥 Koji

Our final question: In 2025, what investment directions will ZhenFund and you personally find most interesting? Especially, are there any non-consensus views, differentiated perspectives you hold?

👦🏻 Yusen Dai

Our differentiated views mainly have three aspects:

First, we're relatively cautious about "time-killing" applications.

Many people are now trying to find the next ByteDance based on ByteDance's experience — seeking a high-engagement, paid-acquisition-driven consumer application. But I think with user attention already so heavily captured by ByteDance, the next killer application may not emerge in this paradigm. That is to say, the next ByteDance probably won't look like ByteDance.

Second, regarding humanoid robots, currently the hottest topic. We've seen many humanoid robot body companies raise substantial funding, but the technical path for general-purpose humanoid robots — whether sim-to-real, training from video, or manipulation data collection — hasn't converged. How to collect such data at scale remains an open question. We believe investment sentiment in this field is overheated. The time required for humanoid robots to complete physical world tasks, let alone enter households for chores, may be far longer than current projections and investment cycles suggest. So we're relatively cautious on the body itself, though we've invested in important upstream components like dexterous hands and motors. In this field, compared to current enthusiasm, we remain relatively sober.

Third, regarding AI applications in productivity. We've observed that in the United States, agents are landing quickly in productivity and enterprise service domains. But in China, because of the widespread belief that enterprises won't pay for tools, there's been much resistance and challenge, and many enterprise service entrepreneurs and investors have been hurt. But I'm thinking — when sentiment becomes extremely one-sided, it often signals opportunity for reversal.

If selling tools alone may be hard to make work in China, delivering the work output itself at one-tenth the cost — that might be something enterprises would actually buy. This could become a powerful AI outsourcing model, not traditional HR outsourcing, but delegating tasks to AI agents to complete.

We're thinking, if we absolutely have to find a To C entertainment use case in AI, that's always going to be pretty difficult. But can we combine the massive breakthrough in productivity with product execution in China? I think enterprise services may not be set in stone — there might be new opportunities here.

🚥 Koji

So what are your key focus areas? What you just shared were some non-consensus views, or directions you think warrant caution and deeper thought. But what are you actively looking at?

👦🏻 Yusen Dai

Since we're always a founder-centric fund, we don't pre-define focus areas every year. But speaking personally, I think having AI do agent-like things in various forms this year will be a very important area. At the same time, I think what I mentioned earlier about achieving scalable personalization through AI programming or modality conversion is also important.

For example, one AI education company we invested in is exploring how to make education sufficiently personalized through AI.

Traditional internet education solved the scalability problem in education — using internet methods to make elite teachers accessible to more people. But going further, we want to achieve personalization while maintaining scale. That's an important opportunity AI brings us.

So we mentioned Yuaiweiwu, this AIGC technology company, and we've also invested in some companies trying to do something like bolt.new — personalized applications generated through AI coding. But this space is still very early stage, and will definitely require lots of adjustment.

🚥 Koji

Let's wrap up here. First episode of 2025, we talked for a long time, and there's a lot of information. But what's most important isn't just the information — it's hoping this episode conveys more optimistic signals and energy, so everyone can take more action, create and produce more. If anyone wants to raise funding, welcome to reach out to ZhenFund. Thanks again, Yusen.

👦🏻 Yusen Dai

Thanks for having me. Your summary just now was really good. When technology waves are this turbulent, although there are still many problems and unrealized ideas, the pace of actual implementation these past two years has far exceeded my expectations. So we have many reasons to stay optimistic and try to break through.

Spend enough time, even spend a little money, to experience the latest AI products, to feel the incremental progress they might bring — this is meaningful and valuable for all of us, whether as investors, founders, or simply as people curious about the future.

And of course, welcome everyone to listen more to "Crossing" and ZhenFund's "This Is Serious" podcast — it will help us learn and understand AI better.

🚥 Koji

Alright, thank you everyone, happy new year, bye.

👦🏻 Yusen Dai

Happy new year everyone.


Subscribe to the "Crossing" Podcast

🚦 We track the industry shifts and new entrepreneurial opportunities brought by the new wave of AI technology. "Crossing" is Steve Jobs's metaphor for Apple — standing at the intersection of technology and liberal arts, where great products are born. AI is transforming every industry. We seek out, interview, and rally "active doers" of the AI era, exploring and embracing new changes and new possibilities together.

👦🏻 Host Koji: Co-founder of The Fair and Tangdao. I believe technology, especially AI, will fundamentally transform society and empower humanity. Welcome to chat with me, bounce ideas, and connect on what's next. Koji on Jike, Koji's website

👧🏻 Host Ronghui: Works at a tech VC, former Silicon Valley correspondent for CBNweekly. Ronghui on Jike

Join the "Crossing" Membership Group

☀️ First-hand AI news and insights

👫🏻 We encourage everyone to date/make friends/find future collaborators

🦀 Add our assistant on WeChat to join: Rwkfbcianvd, or scan the QR code below

References [1] This Is Serious

[2] Crossing

[3] SWE-bench

[4] Koji on Jike

[5] Koji's website

[6] Ronghui on Jike