"After Putting Down the Hammer Looking for Nails, We Made Our First Million Dollars" | A Conversation with Mootion Co-Founder Tong Chao on Product-Making, the AI Video Startup Landscape, and Going Global

Stop looking at the hammer and look at the nail.

From models to applications, AI video products are proliferating.

This week, we invited Tong Chao, co-founder and Chief Product Officer of Mootion — a company that has focused on overseas markets since day one, reached 2 million users by early 2025, and hit seven-figure ARR — to share Mootion's entrepreneurial journey in the AI video space.

You'll hear how Mootion users actually use the product, Tong's reflections on "hammer-looking-for-nail" AI entrepreneurship, his views on whether model improvements will render applications obsolete, his insights on improving product experience, his analysis of the AI video landscape, and Mootion's perspective on Brazil and Arab markets as a company that went global from day one.

Listen on WeChat:

Listen on Xiaoyuzhou:

🚥 Koji:

This week's guest is Tong Chao. Tong Chao is co-founder and CPO of Mootion, an AI video app. Let's start with a quick-fire round. Age?

👦🏻 Tong Chao:

🚥 Koji:

Alma mater?

👦🏻 Tong Chao:

George Washington University and The Hong Kong University of Science and Technology.

🚥 Koji:

Work experience?

👦🏻 Tong Chao:

Currently co-founder of Mootion. Previously AI product lead at 360 and product lead at Innovation Works' AInnovation.

🚥 Koji:

Can you describe Mootion in one sentence?

👦🏻 Tong Chao:

Our positioning for Mootion is: we want to enable people who have never used AI to make their own videos and tell their own stories.

🚥 Koji:

How long since you founded Mootion?

👦🏻 Tong Chao:

Two and a half years.

🚥 Koji:

How much has Mootion raised?

👦🏻 Tong Chao:

Roughly seven to eight million USD so far.

🚥 Koji:

User base?

👦🏻 Tong Chao:

By early 2025, around two million users overseas.

🚥 Koji:

Revenue scale?

👦🏻 Tong Chao:

We also reached roughly seven-figure ARR around the same time at the beginning of the year.

🚥 Koji:

We know there are many AI video products on the market. In two more sentences, introduce yourself and pitch Mootion to everyone?

👦🏻 Tong Chao:

I'm Tong Chao, co-founder of Mootion. Mootion hopes to let people who have never used AI start telling their own stories and creating their own videos. Currently we enable users to complete videos for short-form platforms, personal life scenarios, or industry use cases in just four simple steps.

🚥 Koji:

Have Mootion users created any viral hits? Whether on Xiaohongshu, Douyin, or YouTube?

👦🏻 Tong Chao:

Shortly after launch, I asked our operations colleague to run an experiment — try spending 5 minutes daily making a video with Mootion, and see what kind of response it got on Xiaohongshu.

The interesting phenomenon was, he spent about two months, 5 minutes every day, and now has roughly 13,000 followers. Let me show you some examples from different scenarios. This is a European user, the video is religion-related. He now has roughly a hundred thousand followers, and his recent videos have gotten millions of views. This falls into the short-form video domain. People make a lot of faceless videos and get pretty good viewership.

🚥 Koji:

Faceless means not showing your face?

👦🏻 Tong Chao:

Yes. I also collected some other examples from users. Here's a menu — we specifically have a feature for teachers and students to use AI-generated stories to enhance teaching. For example, this is a bilingual story with both Chinese and English. Additionally, with some of our own technical积累, we also have a self-developed model. We find that in professional video production, there's also quite good usage.

Here's a video produced by a professional video studio in Poland — when this person puts on glasses and enters the virtual environment, all the character footage is generated by us.

🚥 Ronghui:

Is this video a commercial or?

👦🏻 Tong Chao:

It's an MV. They had the song first, and wanted to make a professional 3D video, so they used our model's capabilities. We worked with them to produce this very professional 3D video.

Mootion Use Case Categories: Short-form creators using generation to replace original material editing; animated dialogue-style educational shorts; professional creators producing MVs, commercials, etc.

🚥 Koji:

Podcast listeners can't see the video, and it's hard to switch platforms to search while listening. Please tell us again — what types of videos are users actually making with Mootion?

👦🏻 Tong Chao:

Category one: short-form video creators who completely use generation to replace their original clip-based editing workflow.

Category two: the educational scenario shown in the second video — many teachers and students use Mootion for teaching.

Category three: professional creators who mostly use our underlying capabilities but combine them with their own professional techniques to produce professional content like MVs, video commercials, or even segments in 3D animated shorts.

🚥 Koji:

What's the most viral hit a user has made with Mootion?

👦🏻 Tong Chao:

That's a good question.

Our focus is more on giving users sustained creative ability, rather than continuously producing viral hits ourselves. Internally we've discussed: from a product perspective, our viral hit might have a one-week lifecycle, but after that week, what do we continue doing? Keep chasing viral hits, or more continuously seek longer-term value? Do we serve a one-week viral hit, or do we serve a ten-year user?

We believe our choice is the latter.

🚥 Koji:

Previously, whether Pika, Vozo, or Vigo and other AI video companies, they've all had viral hits that broke the internet. Vozo's founder CY also came on Crossing, mentioning that Vozo's several milestone growth moments were all related to viral breakout hits.

But if Mootion doesn't have such viral hits, does that indirectly suggest that Mootion hasn't truly hit a market pain point?

👦🏻 Tong Chao:

That's a good question, or a good angle to observe from.

From my understanding:

First: video content itself, as a content format, can cover extremely many scenarios — it's not limited to making good content on TikTok or YouTube that can be monetized through multiple channels. Video still has value in many, many domains.

Second: Using viral hits as a growth or operations tactic — I completely agree, it's excellent leverage. In the cold-start phase, or at a new milestone, having very good feedback gets people curious to try your product. I fully agree on this point. From our growth perspective, we may not have done too much on this front. Returning to our thinking just now, I believe doing this more long-term is better.

Third: Whether you've hit user pain points — in our understanding, this shouldn't be observed from a traffic angle, but from a revenue angle. If users are paying you real money, and the payment rate is good, then I think you've undoubtedly hit a pain point. It's just that it may not have gotten good exposure in terms of traffic.

🚥 Koji:

Mootion has been live for some time — what do users who keep renewing their subscriptions typically use Mootion for?

👦🏻 Tong Chao:

It's actually quite typical — in short-form video, that group of people doing faceless content and operating their own accounts are the ones continuously paying.

My view is: AI might get you to 50 points in terms of output, where before you might have only had 20, and now you can get to 50. But to operate an account well well, the remaining 30 or 40 points depend on the user themselves. Do you have a good story? Good expression? Good presentation? Good operations? These things aren't directly related to AI — they're directly related to the user.

Second part, back to teachers and students. I once did a user interview where a teacher told me about an experiment he tried. When explaining a certain knowledge point, he made a two-minute video with Mootion to show students, and found that student engagement in class suddenly became very good.

🚥 Koji:

Recently, there's been an AI video making the rounds — a Tom and Jerry episode. Many viewers didn't realize it was AI-generated at first; they thought it was just another episode from the classic series.

When I saw it, my first thought was that it looked a lot like the kind of video Mootion wants to make. It had a script, a plot, consistent characters, ran about a minute long — all seemingly aligned with Mootion's goals. But we also know that Tom and Jerry video basically emerged from a base model upgrade. There was almost no engineering, almost no human interaction involved.

So I really want to ask you — did it make you anxious? That one day, after you've built so many product features and optimized so many workflows, a single upgrade to AI base model capabilities could just wash it all away, reduce all your effort to nothing?

👦🏻 Tong Chao:

Actually, quite a few friends have been talking to me about that Tom & Jerry thing over the past week or two.

Here's how I think about it. First, from a technical standpoint, test-time training is definitely a promising direction — we're actually doing some follow-up research ourselves. I believe this capability could potentially push video generation to the next stage.

Second, for video applications like Mootion, we actually see base model upgrades as a good thing, not a threat. Why? When we observe how users create, video production is a task with many stages, a long workflow, and many stakeholders involved. It's not like Jasper or other text-based products where the workflow is relatively short — I think, then I output.

We believe that the longer the task, the more stages involved, the more stakeholders there are, the harder it is for AI to disrupt.

It's difficult for AI to design good experiences and good features within such a long process, with so many stages and complex interests at play.

🚥 Koji:

You just mentioned that video production has so many stages and so many roles involved, so it won't be as easily overwhelmed by base model capabilities as text work, because there are so many details. But there's also a counterargument — people feel that a lot of folks building applications today are just building wrappers.

👦🏻 Tong Chao:

If you look at the history of technology development, the word "wrapper" tends to appear during technology hype cycles, and nobody mentions it afterward. I think AI will be the same.

Right now we're in a period of rapid AI development, and many people significantly overestimate where the technology currently stands. When you overestimate it, everything built around it looks like a thin layer. But from our perspective, when you actually build products, you find that base capability is 50 points. But what users want is at least 80 points. Who fills that 30-point gap? The product does.

Mootion is essentially filling that 30-point gap. So for us, whether it's called a wrapper or not isn't really the point. Assuming future capabilities like the Tom and Jerry example you mentioned — after base models improve, we believe application and product design capabilities will become even more differentiated. That differentiation is what drives users to choose A over B, or to choose anything at all.

As entrepreneurs, I think the key lessons are: first, don't have a technology purism complex, only wanting to do cutting-edge things nobody's done before; second, don't have technology hallucinations, thinking the tech is already great or will become invincible in two months. We need to be more practical, have an objective assessment, and focus on what we're doing.

When Model Improvements Become Commodities, Product Polish Is Better Done Early Than Late

🚥 Koji:

Yueguang Zhang made a rare public appearance recently, and someone sent me photos of his presentation slides. One slide said to look for product opportunities in the 90-to-100-point range.

Because users don't accept 20-point, 40-point, 60-point, or 80-point products. So this seems somewhat different from what you just said. Are you worried that today the large model is at 50 points, and you spend a lot of time getting it to 80 — but there are also many people waiting for the large model to jump straight to 90, and only then building products?

👦🏻 Tong Chao:

This is potentially a real concern for us, but looking at the current direction of technology and what it actually takes to meet user demands — that 30-point gap, as I mentioned, video creation has very long stages, many processes, and many stakeholders involved.

There's something extremely important in there: your know-how about video creation and your understanding of users. Even when models reach 90 points, that 90 becomes 60 — it becomes a commodity, the bar rises, it becomes the new 60. Who fills the remaining 30 points? Still product companies. And on this point, doing it early is definitely better than doing it late. As for whether to wait until model capabilities hit 90 before building, I think that's a miscalculation about technology.

🚥 Koji:

Can you give us one or two examples? Maybe the model today is only 50 or 60 points, but you've significantly improved the experience through product design, so users end up with great work — examples in terms of interaction or design?

👦🏻 Tong Chao:

First, I have a personal habit: since our launch last year, I interview at least one user every week, for at least 20 minutes, regardless of country. That's accumulated a lot of user feedback. One consistent theme: users generally consider Mootion an easy-to-use product, very simple.

We tell users that Mootion lets you complete a video in four steps. Each step is like a guide — very simple, foolproof operations. In video creation, from the starting point, how to write scripts, how to express shots, what narrative types to use — connecting all these processes and automating them with AI is how we achieved that 50-to-80-point capability I mentioned. What users see is an 80-point product.

The second thing is also quite interesting. Last December, we held a small user meetup in Tokyo. During the conversation, a man in his sixties, a director, asked me why the platform couldn't generate Asian-looking people when generating Asian characters. Because from the model's perspective, a lot of training data is Western-biased. So even when users specify generating a Japanese or Chinese person, the output easily has Westernized features. He said: "Why can't you do something about this?" This was actually a detail I had overlooked at the time.

When I got back, I specifically worked on model optimization to improve instruction following, so that generating different ethnic appearances would be more accurate. Later I learned who this person was — a very famous Japanese producer and director named Takashi Asai. Zhou Xun's first film, Suzhou River — he was the producer.

Third thing: in February, I traveled across the Middle East. I found that March was their Ramadan, they didn't have much to do, so they needed to create a lot of content. So we quickly launched a new feature to generate Islamic stories well, for them to share.

What was the issue here? In Islamic doctrine, they have Allah and twenty-three prophets. These figures cannot have visual representations in doctrine — they're usually depicted as light or a halo, very sacred. We looked across all AI generation capabilities to see which one could generate such content by default. None could. Mootion launched this feature quickly on March 1st, and volume took off very fast — at one point, this content was driving maybe ten-plus percent of daily usage.

One last thing I find interesting: AI inference costs cannot be ignored today. We want everyone to use it, so we need to make it cheaper, otherwise we can't scale. From this perspective, we've actually invested enormously in backend architecture and inference optimization. This is invisible to users, but what users see is that we're affordable enough — with a very clean experience and extremely low price, they can generate good enough content. Since our product launch in late June last year, in about half a year, we've optimized our gross margin space by roughly 50%.

🚥 Ronghui:

You mentioned four steps to complete a video creation earlier — how did you arrive at four? Why not three, why not five?

You also said you connected the entire video production workflow and automated it — anything in this process that would be instructive for entrepreneurs?

👦🏻 Tong Chao:

The four steps came from deconstructing the entire video creation process.

We identified what users need to do.

First, what they want, what they have.

Second, getting a complete content structure.

Third, after organizing the main functions and content, choosing叠加 or附加 elements — like effects, transitions, sound, and so on.

Fourth, composition and sharing. Composition and sharing means that a single video may not be enough to support content distribution — things like how to write descriptions, how to do hashtags — so we defined it as four steps.

For the automation part, I'm very grateful to one of our investors, a very professional, top-tier domestic film and television company. From a film and television creation perspective — things like how to write truly good script structures, how to express shots — they gave a lot of advice. For example, a frightened expression definitely requires a close-up of the eyes. What expresses a joyful scene? Definitely an ultra-wide angle with lots and lots of people. These are things we need to define as content types that AI can generate, including narrative types.

🚥 Koji:

I'm curious — why wouldn't AI already know that fear should be expressed through close-ups of eyes, and joy through wide angles with lots of happy people? Why do you have to do this yourselves?

Another related question: you mentioned gross margin optimization, which seems to go against advice many people give entrepreneurs today — don't optimize costs, because large models will quickly drive costs down themselves. It sounds like you're doing things somewhat against the current or against mainstream opinion. What do you think?

Tong Chao:

On the first question, we did expect large models to have this capability when we were testing. But the truth is, they don't understand. Even now, they still don't.

On the second point about gross margins — this is actually quite interesting. Our initial goal was to optimize the inference architecture. Because it was too slow. To generate a good video, we wanted it to generate quickly. But we found that slowness was fundamentally about cost, because you were consuming more GPU hours, so you had higher costs. When I make it faster, costs naturally come down.

Koji:

I really find it hard to believe that large models don't know a frightened scene should use a close-up of a person's face. So while you were answering my question just now, I had Claude generate a text-to-video prompt for me, with a clear visual requirement. Those six words I gave it were: "a terrified woman." And it produced an extended text prompt.

In that prompt, it explicitly wrote about a young woman with wide eyes and dilated pupils, with the camera slowly pulling back from a close-up of her face to show her isolated, helpless situation, and so on. So what I want to say is, I feel like models can do this today. Could it be that six months ago they couldn't, and you spent a lot of time training using the methods you just described — do you worry that time and effort was wasted?

Tong Chao:

If you test it in isolation, sometimes it works.

But put it into a complete engineering system where this is a user's prompt and raw material, and you need to organize it into effective, complete content. This story contains 64 scenes, and across these 64 scenes, the semantics need to connect, and every piece of content, every image, every video needs to express the meaning of that moment — that's extremely difficult. So if you only give it one input, it can work. But when you make it a system, make it a network, it fails. That's a huge difference.

Koji:

I'm quite curious — how many R&D people do you have at your company now, and what's the composition?

Tong Chao:

Currently very few, around 20 people. We have just one finance person, one operations person, and everyone else is R&D. Of those, roughly half are algorithms, the other half engineering.

Ronghui:

Were there any particular inflection points where revenue clearly correlated with feature releases?

Tong Chao:

In December last year, we launched a very important feature. It might sound simple: templates.

Templates are essentially different entry points.

These templates are different from design templates — behind them are AI workflows. This actually gets back to what Koji was asking about: why does it seem like models can do this, yet large models can't do it themselves?

Behind it all, you need many different workflows. For example, that Islamic story I mentioned earlier — behind it is a long chain of AI workflows, or you could understand it as an agent working, completing a series of tasks to produce good content.

So I think this was a very important change. Through this kind of interface, we were able to go deeper and faster into user scenarios. We found that when users got good content through these different entry points, there was a significant increase in paid conversions.

Koji:

While using Mootion, I noticed you have a design where users can't choose the video model themselves. So users presumably can't freely select whether to use Flux, Keling AI, Veo2, and so on. What was the thinking behind this design?

Tong Chao:

This relates to our positioning.

As I mentioned, we want people who have never used AI to be able to create their own content and tell their own stories. So we believe users want results, not models. As Zhao Benshan said, don't watch the ads, watch the cure. They don't care what this is or what brand is behind it — they want the final result. So we make our own selections; giving users model choices actually increases their learning costs.

They'd ask: What is Keling AI? What is Veo? What is Hailuo AI? What is Pika? What is Sora? We're all in the so-called AI circle, so we absorb this information easily. But AI penetration is far from the point where everyone knows these things. There are too many people without this information, and when they don't have it, I think not giving it to them — letting them focus on results and creation — is better.

Koji:

I'm thinking, as you just mentioned, the day before we recorded this podcast, Keling AI released 2.0 and it was genuinely impressive. The Crossing WeChat public account published a long article introducing Keling AI 2.0. Surely many users saw this news to some degree. Would they come ask Mootion: do you have Keling AI 2.0 inside? Or another scenario: would you lose users who go use Keling AI 2.0 directly, or go use competing products similar to Mootion that support Keling AI 2.0?

How do you view this? It sounds like you're just giving up on these users.

Tong Chao:

Yes, because I actually understand these users quite well. We can broadly call them AI early adopters, or technology early adopters — they're like the first 3% on the adoption curve.

These people are very active, their opinions can influence others, so we see an enormous amount of messaging from them. This group has very strong appetite for new things. But frankly, these users have very short lifecycles — roughly one month. After a month, given current technology iteration speed, there will definitely be something new for them to use. They don't care about using something for a specific scenario, because for them the greatest value is "freshness," not any particular use case. Switching to the next "fresh" thing has very low switching cost for them.

For us, I think it's a tradeoff. Building products is often about tradeoffs. Immature technology, insufficient resources, unclear users — a whole series of things. I think tradeoffs are crucial, because you can't get all resources, and you can't have sufficient voice at any given moment. So making good tradeoffs, and being able to identify the most important problem — that should be the core of product work.

Startup Lessons: Founders Should Keep Doing User Interviews, Avoid "Looking for Nails With a Hammer in Hand"

Ronghui:

You mentioned you do a user interview every week — that's quite interesting. Most founders will say they care deeply about user feedback, but doing an interview every week and maintaining that frequency is still quite challenging in terms of execution and persistence. Are you still doing it?

Tong Chao:

I'm still doing it. I send emails to about 50 users every week — completely cold outreach. Maybe one or two people respond.

For users, the response rate through email communication is very low. But I think this is extremely important, because we've talked about a lot of AI technology. When this technology is immature, user mindset and user acceptance points are very tricky. Where exactly they are is very hard to find — you only know by asking them. But you have many types of users, so you need to constantly find them. That's why I persist with this myself.

Also, I think as a product lead, having this kind of intuition is important. When you're building a product yourself, it's easy to be prejudiced. And when technology is also changing, it's easy to lose that intuition, to not know what users are actually thinking. And when their acceptance point is shifting, you need to know at all times.

How do you achieve this? I think in our current state, it's only through constant communication with users.

Ronghui:

What you're saying reminds me of a tweet Stripe CEO Patrick once posted. The gist was that he didn't treat user research as user research pointing to what features or products to build. Rather, it first lets you form your own mental model, and then you use that mental model to build a product.

Tong Chao:

Exactly, I strongly agree. It's a bit like what I mentioned about intuition. For so-called high-cognition people, it's actually quite easy to lose user consciousness.

Because your own cognition is high, you tend to think what you're saying might be right. But the people you're serving might be a group you completely don't understand — so how can you empathize with them, find intersection with them? I think you have to constantly throw yourself into their scenarios, into their situations, to get the direction for moving forward. I really agree with what was just said.

Second, you can quickly switch to another identity, or you can have two or three or five identities simultaneously to think about: how should I build my product?

Ronghui:

Did this come from some previous lesson or experience?

Tong Chao:

I think so.

I actually studied CS, then learned Machine Learning — I have a technical background. Just speaking for myself, people with technical backgrounds are relatively prone to technical hallucination. You think technology can do everything, but when you get your hands "dirty," you find it can't, and you get stuck there. So my past experiences gave me similar negative feedback. I have to constantly remind myself to throw away the technology.

Whatever the technology, when I'm trying to solve a user problem — what exactly is the problem? I constantly need to remind myself, or in past product experiences, I've made many attempts where we were looking for nails with a hammer in hand. Sometimes it succeeds, because your hammer happens to fit that scenario and hits it just right. But mostly it fails. So these failure experiences were very strong negative feedback for me, especially for this current product.

Koji:

So was there any specific story of looking for nails with a hammer in hand, not finding them, that made you feel particularly frustrated?

Tong Chao:

Too many to count. From Mootion's perspective, there's one that's quite interesting — I wouldn't call it a frustration, exactly, but it was an important lesson. In the beginning, we were very focused on our foundational R&D. We built a 3D motion generation model, not large in scale, but probably the world's first. Its purpose was to generate character movements like those you see in 3D films, animation, and games. You input a text prompt, it outputs a motion. Around mid-2023, we got this model and shipped it. We quickly received positive feedback from over twenty thousand professional 3D creators.

Two months later, I spotted a problem — the one I mentioned earlier about what comes next. What was this model actually for? Generate a character animation, and then what? Could it be used in actual game or film production? Not yet. And if not, where could this go? Nowhere.

So we made a quick decision: this wasn't going to work. Two options. First, build a highly specialized workflow to serve professional creators. Second, kill it and do something else. We decided very fast to internalize this model's capabilities into Mootion's current product, to support the controllability of our generation.

So when I said earlier that I constantly need to remind myself — it's because I have this tendency too. With so many technical possibilities and opportunities in front of you, you can't help reaching out to explore. But when you go out without thinking through what problem you're solving, you'll definitely get pushed back. This really matters. That was four to five months of entrepreneurial effort, and the cost was quite high.

🚥 Ronghui:

Speaking of which, does this connect to your previous work experience? You mentioned 360 Mobile and AInnovation earlier. Could you talk about how those experiences relate, and what key takeaways you drew from them?

👦🏻 Tong Chao:

Yes, I think AInnovation is one example. It was essentially a previous-generation AI company, roughly from 2016 to 2022, part of that era including the "Four Little Dragons."

That generation of companies was deeply influenced by DeepMind. DeepMind gave everyone tremendous hope — beyond chess, it could play the far more complex game of Go, making excellent progress across countless domains. So AInnovation, and that wave of AI companies, all started with a kind of AI technology dream. But the problem was exactly there: if you're holding technology, you're the classic case of looking for nails with a hammer in hand.

There were two issues. First, the hammer wasn't hard enough at the time. CV-centric technology, or NLP, or machine learning — in certain niche scenarios, you might achieve performance suitable for industrial or commercial use, but generalizability was extremely weak. So even when your hammer hit the nail, you needed huge amounts of people and additional engineering to supplement or deliver the solution. It was a very heavy model, not what we imagined AI could do to improve or even multiply labor productivity.

The second problem: still looking for nails with a hammer. As a startup, even with substantial funding, how many chances do you get to swing that hammer? They're limited. The more you swing, the higher the risk becomes. Your willingness to keep swinging shrinks, because you don't know if it'll work. From this angle, if you haven't hit enough nails early on, and later can't sustainably drive revenue, you face enormous challenges. Experiences like this are exactly why I must stay vigilant — even when using new technology to build products or applications, I must first think clearly about who our users are and what problem we're actually solving.

🚥 Koji:

AInnovation was a very high-profile company at the time — Kai-Fu Lee went all-in on it. There was news about the CEO following Kai-Fu and giving up an 18 million RMB annual salary, seven funding rounds, with international giants like SoftBank behind it. Such a golden lineup, and they worked hard to list on the Hong Kong stock exchange. But after IPO, massive losses year after year, with the stock price now down 80%. As Product Director, you must have been among the core executives. You mentioned how painful it was, looking for nails with a hammer.

Could you talk about what nails you tried to find back then? Stories from that era might offer some lessons for what's happening now.

👦🏻 Tong Chao:

I very much agree with your point — even with generational shifts in technology, the scenarios people face are quite similar.

You can see this with SenseTime too; they're going through a transformation. Returning to the scenarios: we actually explored many industries — retail, industrial, manufacturing — which were very large or high-volume sectors, each with its own problems.

The problem with industry is that it has extremely fragmented sub-sectors. Chinese manufacturing was growing incredibly fast, which meant their requirements were getting higher and higher. Across different sub-sectors, these high demands exposed AI's limitations at the time. This was partly constrained by technology — when you swung the hammer, you couldn't necessarily hit the nail.

Financial services. I think finance was genuinely one of the most promising industries for B2B, because people in finance are well-educated and open to new technologies. It could also be an excellent revenue source. But finance had its own problem back then — very strong demands around model ownership and ongoing maintenance. You needed to deliver a model, and its ownership, subsequent updates, and maintenance had to belong to the bank, or insurer, or securities firm. They all had similar requirements. This was quite difficult for the previous generation of AI companies. You deployed your model, then had to continuously update and maintain it. Unlike today where OpenAI can lock a version and solve many problems — back then you couldn't do that. You had to keep at it, which meant committing more human resources to maintenance and updates. That was a very challenging problem.

Marketing was too small. Every scenario was niche. Plenty of opportunities, sure, but when each scenario is small, and you don't have the generalizability that today's AI technology offers, the challenge becomes: how many horizontal people do you need to deploy? Ten, fifty, a hundred projects or client accounts? That was a difficult problem.

The defining feature of this AI wave is universal access — we can see many possible evolutionary paths for technology

🚥 Ronghui:

Do you see any similar cases from your AInnovation experience in this current AI wave?

👦🏻 Tong Chao:

For this wave of AI, my assessment is probably not.

Here's an example. Eight or nine years ago, it would have been hard to imagine AI enabling an elderly grandmother in a county-level town to talk with AI. But now in China, I'm sure you can find such cases. This past Spring Festival when I went home, a distant elderly relative said to me: "There's something called Seek, seems pretty good, help me find it." He meant DeepSeek.

The reason behind this, I believe, is that compared to before, the biggest difference in AI's capabilities this time is universal access. It enables more ordinary people, those who've never used AI, to benefit from AI's capabilities. This is what I consider the greatest, greatest value of this generation of GenAI. Actually around 2022, we built a Transformer-architecture model — small, but that's not the point. The greatest value of this generation is universal access. When you have universal access, your hammer now has better capabilities to actually hit the nails.

From this perspective, it's completely incomparable to the previous generation's approach. So I believe this generation of AI, whether in applications or models, probably has enormous potential, and likely won't face the collective flame-out that happened to the previous generation.

That said, if we look only at large model companies, I think this question still hangs over them. You've built a model, then innovators like DeepSeek emerge, creating significant shocks. How do you face the market, how do you face your own technology? At this point, I do think large model companies may still face this problem. But for those doing applications, this burden is much lighter.

🚥 Ronghui:

In a previous conversation, you mentioned meeting Ilya at OpenAI.

👦🏻 Tong Chao:

Yes, it was quite coincidental. We'd just started our company not long before, and Kai-Fu had introduced my co-founder to go to Silicon Valley for about a month. By chance, roughly a week before ChatGPT launched, they met with Ilya and Greg. The message my co-founder brought back was that they seemed to be building an application, but Ilya himself said he didn't know how it would go. And then it happened — ChatGPT launched and became a world-changing application.

What struck me about this, beyond the hammer-looking-for-nails lesson, was a second point: when technology is evolving rapidly and innovative products keep emerging, no one has superior cognition to anyone else, especially when facing the market. Everyone is more or less the same — all exploring, enjoying the fruits of technology, searching for potential value in the market or on the user side.

So I think the greatest inspiration for us is this: when we're in this AI era, the most important thing, in my view, is to practice extensively, fail fast, and iterate at full capacity.

Actually, Ilya and Greg didn't share much so-called inside information or deep insights, but their expectations and the subsequent reversal gave us tremendous inspiration.

A panoramic view of AI video entrepreneurship: analyzing five types of companies one by one

🚥 Ronghui:

A long time ago, I heard a talk by Steve Chen, one of YouTube's founders. He said something that left a deep impression. Many people assume a great entrepreneur is like a god, a prophet who knows something and then goes and does it. Chen said that's not true at all — it's the biggest misconception about founders.

Ronghui:

Coming back to AI video entrepreneurship — could Tong Chao help us map out the full landscape of AI video startups, and maybe highlight which ones you're most bullish on?

Tong Chao:

Model companies are fairly evenly matched — everyone's competing on the same playing field to push performance. From the application side, I roughly divide video into five categories.

The first is AI-assisted editing — using AI to do what traditional editing did, companies like Descript, Opus Clip, and Captions.

The second is talking head — there's a person in it, either a digital human or a real person turned into one. The typical example is HeyGen. I also saw this morning that Synthesia, a UK company, just crossed $100 million ARR. This category ultimately produces video too, but the path involves having a real person in it, used across different scenarios.

The third is video effects — Pika or Vigo, for example. What they do is create different effects inside a video. It's a lot like the post-production effects in traditional filmmaking, but they solve it through generation.

The fourth is more like Mootion — complete end-to-end, generating full content rather than a video clip. Video model companies are more like generating a clip, but these application-layer products generate complete content. Mootion, Flicky, and a US company called Invideo — they're all roughly in this space.

The last category is single-point features with different functions — VideoDubbing, video translation, or Face Swap. These are more like specialized tools for specific steps in the video creation process.

That's roughly how I categorize video applications. But I've noticed something: all these companies are shovels — they get you part of a final piece of content, or create one link in the chain. Above these shovels, I think generation and editing will eventually merge.

From the user's perspective, I don't care whether a clip or effect was generated or shot, or how I cut a segment to get, say, a highlights video.

We think a likely future scenario is that AI editing, or traditional editing, and generation — whether you're doing effects or full-content generation like Mootion — will probably converge.

On top of this convergence trend, there's an opportunity: whoever first defines what the product looks like when editing and generation fuse. I think this is a massive opportunity. A little preview: Mootion may launch the first product in this form around mid-year or the second half of this year. Stay tuned — it's the crystallization of our thinking.

Second, Ronghui asked which products I'm watching. For product innovation and experience, I really admire Descript. Descript is trying to use natural language interaction to update or iterate on the cumbersome capabilities of current video editing. I use it daily — there are many small issues. But the direction is a great signal of what role AI might play in editing and video creation.

Another product fewer people may know is also from a UK company, called Veed.io. The product is easy to understand, but what's interesting is that it's a product that grows with its users. Veed first launched around 2019 with almost no users. The founder was completely immersed in the same environment as his users. He openly shared his entire entrepreneurial journey on Twitter and other social media. So tons of users gave them feedback, interacted, engaged.

They built many features, and around last year or the year before, started integrating AI and upgrading themselves. The very first "VideoGPT" in ChatGPT was built by Veed. So it's a very agile team and product. They hit $1 million ARR in about two years, then another three years to what I think is now over $20 million ARR — very fast growth. But behind it, you see their iteration and direction responding to user needs in real time.

Koji:

There are so many Chinese entrepreneurs in this generation of AI video founders. Who do you admire most?

Tong Chao:

Joshua from HeyGen — I really admire him, worth learning from. He's a bit like that Veed founder I mentioned. I know their product went through pivots, but they did a lot of user-friendly designs and iterations to follow the market. I think this may be a very important trait, or advantage, of Chinese entrepreneurs.

Koji:

Is there anyone building their own AI video product — not yet launched, or launched but not getting enough market attention — who you believe will achieve something remarkable given time?

Tong Chao:

I'm hoping to see more foundational video models emerge domestically. Right now Keling AI is probably the furthest ahead, but there should be different foundational approaches to getting better video results. From this angle, I'm quite looking forward to Cao Yue's Sand AI — I heard they may have some releases soon.

Koji:

Have you worked with him directly?

Tong Chao:

Not directly, but since we're both Sinovation Ventures portfolio companies, there's some indirect connection.

Koji:

What are you excited about?

Tong Chao:

They're using a different technical approach. The details aren't public, but different technical routes may yield different good results — possibly diverging from DiT.

Koji:

Being on the front lines, you must feel intensely how from Sora onward, all these video models chasing each other — within one year it feels like ten years of evolution. Which team impresses you most?

Tong Chao:

In large video models, I think the Keling AI team is genuinely impressive. I once attended one of their version launches and had some exchanges with their team. After Sora, Keling AI actually went quiet for a while — Sora happened to land right in the middle of their R&D cycle.

But their results, including after Keling AI went overseas, and the feedback from domestic and international creators — I find that quite rare. The difference here is between a content company and what we understand as an internet company. They reached SOTA in a very short time, then converted that SOTA into a foundational model product that massive numbers of users love and keep using.

Koji:

In your observation, why could they do so well? Of course Kuaishou has money and user data — everyone knows that. Is there any "secret sauce" you've observed or sensed behind their success?

Tong Chao:

I think it's persistence. For Keling AI, more of the results probably came from the research team.

The product layer isn't that heavy. So on the research side, I think they were quite persistent — committed to a technical route, made significant investment. They didn't panic after Sora came out. I think this is important: being able to steadily produce good content, good results, then productize and update at their own pace. This is, at least up to now, a relatively important attribute or differentiator domestically.

Koji:

To your knowledge, who are the soul figures inside Keling AI that were key to their success?

Tong Chao:

If you ask me to name names suddenly, I might not remember. But there are indeed two to three quite critical people guiding the entire research direction.

Ronghui:

Or were there strategically correct decisions? Does persistence yield good results only if you chose the right path to begin with?

Tong Chao:

I think they bet correctly on the technical route — the Sora-like approach. This was definitely right; if they'd been wrong, they probably couldn't have delivered such results in so short a time.

Second, on strategy — this may not be the product team's credit — I think they very quickly went international, went global, extended their product overseas. I'd even rate this as possibly better than Kuaishou's own overseas expansion.

Koji:

Now with giants being so aggressive — whether it's Kuaishou's Keling AI or ByteDance. Today for startups also building video models, like DeepVision or PixVerse, do you think they still have hope?

Tong Chao:

I think foundational video model capabilities haven't converged yet — they're still in a divergent phase, which may be a big difference from language models. In this state, I think entrepreneurial teams with new ideas, new directions, new results and giants' sustained high-investment new results may run in parallel for a while.

But the ultimate outcome — I can't say. We couldn't have predicted DeepSeek suddenly emerging during Spring Festival and then everything converging. I think the video space may run a bit longer, but who ultimately breaks through? I think everyone has a chance.

Koji:

There's a saying today that wherever ByteDance goes, nothing survives. Especially in video — whether it's CapCut, or Pippit AI, which topped Product Hunt last week, another ByteDance video tool.

Of all the various AI video products ByteDance has built, which do you like most? Which are you most bullish on?

Tong Chao:

I think Dreamina is a pretty good foundation. Dreamina has made some very bold product attempts. You'll notice when you open Dreamina, the first screen is a Douyin-like content feed. It's genuinely trying to turn AI content into something users can spend time on, continuously consume — and that matters.

This means Dreamina could potentially evolve from a pure productivity tool into one where the content produced by that tool can be well-consumed by users. That's a very bold attempt. The person leading Dreamina is also a friend I know. But this is still an experiment for now — whether it sticks, or whether it can become a new Douyin, is a possibility, but there's a long road ahead.

Mootion's Three Largest Markets: Brazil, the Arab World, and the United States

🚥 Ronghui:

Let's wrap up by talking about going global. We know Mootion has always focused on overseas markets. Tong Chao, could you tell us where your three largest user bases are right now and why you think that is?

👦🏻 Tong Chao:

By broad region: first is Brazil. Brazil is our biggest market. Second is primarily the Arab world — the six Gulf countries of the Middle East. Third is the United States.

From what we've observed, Brazil was actually within our expectations. Brazil is essentially the first stop for many Silicon Valley startups going global right now. They're also very bullish on Brazilian users' receptiveness to products and willingness to try new things. And Brazil is a large, single market with absolute scale.

The Arab world was relatively unexpected for us. One very typical country is Oman. I checked the data yesterday — we probably have over 70,000 users in Oman now. Oman has a population of maybe less than 5 million. So our penetration rate in that country is actually very high.

Under that penetration rate, things also happened to align well. In February, Oman's Investment Authority and Ministry of Education invited me over for some exchanges. They were curious why a product like ours could achieve such high penetration. I visited most of Oman's good private and public schools and genuinely saw teachers and students using our product to empower, or rather integrate into, their classrooms and teaching. So that was quite a pleasant surprise for me. Oman actually represents a similar user scenario across the broader Arab region, so this part was unexpected for us.

The United States feels very reasonable to us. North America is an important market when going global.

🚥 Ronghui:

Could you talk about the characteristics of users in these markets?

👦🏻 Tong Chao:

Brazil is a market where users generally accept new things very quickly and are willing to try them. Users themselves organize small user groups to discuss new products or new things. So penetration or user acquisition in Brazil should be quite fast. But Brazilian users are actually quite similar to Chinese users five or ten years ago — their payment habits aren't that great. You need to consider this: when doing cold start, Brazil might be a great market for user acquisition, but you also need to think about how you'll handle payment and conversion three months down the line.

Arab users are a completely different group in terms of user habits. Because they have very serious religious constraints. Their rigor means they're relatively cautious about accepting or converting to new things. But conversely, it's precisely because of their stricter doctrines that these people have very high expectations for new things, so-called AI-powered products. So when you find a small entry point in an Arab country, it's very similar to our experience. When you find that small entry point, it greatly amplifies these people's curiosity about new products or new concepts, so growth can be very fast. For us in Arab countries, this kind of growth is probably 90% organic. We did content with 2-3 influencers; besides that, it's all organic traffic. So once you find the entry point, I think natural word-of-mouth in Arab countries can be very good data. Payment in Arab countries is a bit challenging, but if you choose well among the six Middle Eastern countries, it's much better. The US situation I think everyone is relatively familiar with.

🚥 Ronghui:

You mentioned Arab countries are more cautious — does that mean their loyalty is also higher?

👦🏻 Tong Chao:

Exactly. From our own data, we've found that once Arab users accept and convert, their lifecycle tends to be quite long.

🚥 Ronghui:

You also mentioned doing events in Japan, and in our previous conversations we talked about Taiwan — could you share about these two?

👦🏻 Tong Chao:

Japan and Taiwan are regions where we started seeing good user growth this year. The signal is very clear. In Japan and Taiwan we're also basically growing organically. After reaching maybe tens of thousands of users through organic growth, we do some boosters — for example, we work with influencers, then use their content to influence more users.

But Japan and Taiwan, I think these two markets are very similar. They're both markets where user acceptance is difficult, but once you get in, their loyalty is extremely high and payment rates are extremely high.

🚥 Ronghui:

You mentioned five markets. Have you seen competitors in these places?

👦🏻 Tong Chao:

Actually quite a few. The Arab region has fewer — I think people are indeed relatively less familiar with the Arab region. Brazil, many of my friends in Silicon Valley who are starting companies, their first stop is Brazil. The US has even more. Japan too. Taiwan maybe relatively fewer. So the chances of running into friends are very high.

This time in Oman, a teacher told me what apps he uses. He uses DeepSeek. He showed me PixVerse. I thought that was pretty great — Chinese products have also penetrated to users all around the world.

🚥 Ronghui:

What do you think are the advantages of Chinese entrepreneurs at this point in time?

👦🏻 Tong Chao:

Three advantages.

First, Chinese entrepreneurs, at least many friends I know, are very grounded. Everyone is able to find the path to PMF through many methods even when the technology isn't mature yet, thereby enabling rapid product growth. I think this is very strong. I believe in this regard, Chinese entrepreneurs should be better than many American entrepreneurs. This is an important difference.

Second, I have to say that entrepreneurs who grew up in the super cutthroat market of China's mobile and internet space are essentially doing dimensional reduction when it comes to growth and operations strategy. So especially during a new product's cold start and scaling growth, Chinese entrepreneurs probably have some secret menus in their hands.

Third, in the Chinese entrepreneur environment, there are simultaneously two types of talent. First, very good researchers, even very young researchers. They may not necessarily be top-tier, but there are many young, idea-driven researchers. Second, and very importantly for entrepreneurs, we have very good full-stack engineering teams to coordinate with them. This means we can very quickly pass new technology to users. I think this kind of technology supply isn't easily found in other countries. Put these three together and you'll find Chinese entrepreneurs, especially in AI circles, are able to make rapid breakthroughs very quickly.

🚥 Ronghui:

Last question: Do you think companies going global still need to find a local representative in this era?

👦🏻 Tong Chao:

I believe in AI. So we have no intention of setting up a representative in any country or region.

First, believe in the power of AI. AI should serve as an employee or an intern that can represent us in communicating with users. We're actually already doing this.

Second, believe in the power of users. In this globalized context, your local users are your best representatives. Users themselves have sufficient motivation and energy to help you with expansion, feedback, iteration, and so on in these countries.

🚥 Koji:

Thank you so much for your time today, Tong Chao. We shared Mootion's story and got a lot of education and commentary on the broader AI video space. When we look back at this episode a year from now, we should be able to see a lot of fresh insights and clues. Thank you again, Tong Chao, and you're welcome back at Crossing anytime. Bye bye.

👦🏻 Tong Chao:

Bye bye, thanks everyone.

Subscribe to the "Crossing" Podcast

🚦 Crossing is Steve Jobs' metaphor for Apple — standing at the intersection of technology and liberal arts, where great products are born. AI is bringing change to every industry. We seek out, interview, and bring together a new generation of AI entrepreneurs and active participants in the AI era. Together with them, we explore and embrace new changes, new possibilities.

👦🏻 Host Koji: I co-founded Jiepang / The Fair / Tangdao, and started AI Hacker House, a community space for a new generation of AI entrepreneurs. I believe technology, especially AI, is the greatest value creation opportunity of our generation. Feel free to find me to chat, bounce ideas, and connect on what's next. Koji's Jike[1], Koji's website[2]

👧🏻 Host Ronghui: I've worked at a dollar-denominated VC and spent five years as a Silicon Valley correspondent, following tech development and business stories. Feel free to find me to chat and exchange ideas. Ronghui's Jike[3]

References

[1] Koji's Jike: https://okjk.co/0JSUes

[2] Koji's website: https://koji.super.site/

[3] Ronghui's Jike: https://okjk.co/0cbnYV