2026: The "Big Year" for AI Video Is Coming | A Conversation with OiiOii Founder Naonao: After Stints at WeChat and ByteDance, How to Catch the Next Wave?

"LLMs are supermarkets; Agents are restaurants"

Large models are supermarkets, Agents are restaurants

👦🏻 Podcast interview: Koji

🥷 Edited by: Crossing

🧑‍🎨 Layout: NCon

🚥 I believe 2026 will be a breakout year for AI + video, and many new unicorns will emerge from this space.

So the first AI founder guest on the Crossing podcast in 2026 is Naonao, founder of OiiOii. She works in AI video, and her product attracted massive attention and praise right from its beta launch.

Naonao's background is uniquely compelling — she's one of the few product people in China who has fought in the core business lines of both WeChat and ByteDance. She worked on QQ Mail, led CapCut and Douyin effects, and was once head of Bilibili's animation business. She brings both WeChat's deep insight into human nature and product values, and ByteDance's systematic, data-driven growth methodology. More importantly, she's a creator who has harbored an animation dream for over a decade.

In 2024, she finally started doing what she'd wanted to do since college — using AI to let ordinary people make animation.

In this episode, we cover a lot:

On product: OiiOii rebuilt its architecture four times in two months, from "first-frame-last-frame" to fully embracing Sora 2 — what happened in between? Why does she believe Agents won't be swallowed by large models, but will instead flourish together? On methodology: She uses the "supermarket and restaurant" metaphor to explain the relationship between Agent companies and model companies; uses "left brain and right brain" to summarize the core differences of building products at WeChat versus ByteDance; and shares the three key capabilities for being a good product manager. On entrepreneurship: Why does she say animation is one of the few industries in the business world that "rewards purity and passion"? Why does she believe "conflict is the power to get things done"? On product management: What are the most important things to become an excellent product manager?

If you're an AI founder, product manager, or investor interested in video generation, this episode will give you plenty of insights.

Listen on WeChat:

Listen on Xiaoyuzhou:

🎬 The video podcast is also available on Koji's WeChat Channels, Xiaohongshu, Bilibili, YouTube, and other platforms

Since the full interview is quite long (21,432 characters), here's the table of contents:

🟢 Quick-Fire Q&A

Age, alma mater, MBTI and zodiac sign, one-sentence introduction to current company and product, funding status, revenue and profit, team size, pre-entrepreneurship experience

🟢 The Agent Moment for AI Video: Why Now?

When Sora 2.0 demonstrated astonishing storyboarding capabilities, are startups eaten by "end-to-end" solutions, or entering their best era?

  • Why is Agent the optimal form for animation, rather than traditional GUI tools?
  • The two schools of video generation: the "first-frame-last-frame" school pursuing extreme stability, versus the "reference image + text-to-video" school pursuing cinematic language.
  • Only Agent architecture can snap different models together like Lego bricks.
  • We didn't even expect that parents would use OiiOii daily to make Christmas song MVs for their kids, or animate their pets.

🟢 Why We're Not Afraid of Sora

Large models are supermarkets, Agents are Sichuan restaurants

  • As model vendors grow stronger, where exactly is the moat for application-layer companies?
  • Even when Sora reaches 4.0/5.0, Agent products won't die out — they'll become even more prosperous.
  • The Supermarket and Restaurant Theory: Large models are like supermarkets (providing raw ingredients), while Agents are Sichuan restaurants, Cantonese restaurants (providing finished dishes with specific flavors). You can buy groceries at the supermarket and cook yourself, but restaurants will always have their place.
  • About 60-70% of our work is like "seasoning" in the back kitchen: how to turn those raw model outputs into dishes that suit specific tastes (MVs, educational content, animated dramas).

🟢 Not Just "Video Cursor": Where's the Incremental Market?

  • Animated dramas, UGC communities, self-media creators... who are the real paying customers for this wave of technology, based on our beta findings?
  • Thoughts on CapCut: AI video Agents won't replace CapCut — they'll create incremental value. They'll eat the complex "effects production"环节, but lightweight "editing/trimming/sequencing" still requires traditional timeline tools.
  • Why not build a UGC community directly?
  • A counterintuitive finding: the people who need AI animation tools most may not be creating for mass audiences, but for maintaining social relationships (for children, teachers, partners).

🟢 WeChat's Right Brain, ByteDance's Left Brain: Two Cultivations for Product Managers

Naonao is a rare core PM who has worked in both of China's top product systems, WeChat and Douyin. What similarities and differences did she see?

  • Tencent WeChat (right brain): It's emotional, intuitive. Allen Zhang taught us to read 1,000 pieces of user feedback daily, to train that product intuition of "spotting real needs at a glance."
  • ByteDance (left brain): It's rational, data-driven. Here I learned what "strategy product" means — data isn't just cold numbers, it's the probability of user behavior, telling you how to create trends.
  • The commonality: both have pushed their innate strengths to the extreme. WeChat has taken "human nature" to the extreme; ByteDance has taken "recommendation engines" to the extreme.
  • What are the three most important things for being a good product manager?

🟢 Conflict Is the Power to Get Things Done

  • Why "allow conflict, even create conflict" within the team?
  • Healthy conflict is the best filter for "people who get things done." When everyone is "mission-first," conflict turns into back-to-back camaraderie.
  • Once fired my best friend, got cursed out brutally, but a year later they messaged: "I finally understand you."
  • Was once a rebellious rock-and-roll girl, now has become very peaceful: because I realized shouting outward isn't freedom — true freedom is inner vastness.

🟢 Predictions for 2026

  • A contrarian take on the future: technology's "greater interactivity/higher editing freedom" may actually make products more niche, because the masses are accustomed to "passive consumption."
  • Video models won't achieve "unified dominance" in the near term: because each model company's data annotation standards differ, this leaves huge room for "combinatorial innovation" for Agents.
  • I don't want to express anything myself; I hope OiiOii, like me, is a container. Let those chosen by the "god of animation" tell their stories through this vessel.

🟢 10 "I am ___" Sentences

  • I am a...

👦🏻 Koji

Hello everyone, I'm Koji. This week's guest on Crossing is Naonao.

👩🏻 Naonao

Hello, I'm Naonao.

👦🏻 Koji

Naonao is the founder and CEO of OiiOii, a video Agent. Recently on Jike I saw a user say they hadn't discovered any new product that truly surprised them in the past six months, until they encountered OiiOii. I also tried using OiiOii to make a Christmas song MV for my daughter — she loved it and has been playing it on loop.

👩🏻 Naonao

We do see some parents making little videos for their kids every day now, and teachers using animation to make educational videos for children. This is quite different from the creators we initially envisioned.

🟢 Quick-Fire Q&A

👦🏻 Koji

As per our show's tradition, let's start with quick-fire questions. What's your alma mater?

👩🏻 Naonao

Sun Yat-sen University.

👦🏻 Koji

Your MBTI and zodiac sign?

👩🏻 Naonao

INTJ, Leo.

👦🏻 Koji

Describe what you're doing now in one sentence.

👩🏻 Naonao

Using AI to make animation.

👦🏻 Koji

Current funding status?

👩🏻 Naonao

Raising a Pre-A round, almost closed.

👦🏻 Koji

Can you share current revenue and profit?

👩🏻 Naonao

We've only been around for about 4 months, no revenue or profit yet.

👦🏻 Koji

The product launched just a month ago, right?

👩🏻 Naonao

Yes, the beta launched a month ago. Still invite-only for now.

👦🏻 Koji

Current team size?

👩🏻 Naonao

18-19 full-time.

👦🏻 Koji

What did you do before entrepreneurship?

👩🏻 Naonao

I've always been a product manager, entrepreneur, and worked in video creation — always in video content creation. Including running my own studio, and building creation tools.

👦🏻 Koji

Where did you work as a product manager before?

👩🏻 Naonao

Started at Tencent, in the WeChat Business Group, working on mobile QQ Mail. Started my first company in 2014, building an extreme sports content community — we filmed extreme sports athletes across the country ourselves, accumulated lots of content production experience, ran it for about 6 years.

Later, because I had experience in both content production and product, I joined ByteDance as the product lead for CapCut, and also oversaw special effects for TikTok. And because I love animation so much, I worked on animation-related projects at Bilibili too.

👦🏻 Koji

That's quite a rich background. If you weren't starting a company, what do you think you'd be doing right now?

👩🏻 Naonao

I'd probably be living in seclusion somewhere remote, raising small animals, living a simple life.

👦🏻 Koji

Are you someone who can handle doing nothing?

👩🏻 Naonao

I am now.

👦🏻 Koji

You weren't before?

👩🏻 Naonao

Right. Being idle now is a state that really recharges me — it's about enjoying the present moment. Before, my restlessness came from feeling like I should be doing something when idle. That was a very outward-reaching state.

👦🏻 Koji

What do you mean by "outward-reaching"?

👩🏻 Naonao

Like you can't settle down. Deep down you're thinking, "I have to accomplish something to be worthy."

👦🏻 Koji

Right, you have to construct some meaning.

👩🏻 Naonao

Exactly. But I hadn't actually experienced the meaning of quiet. Once you understand that, quiet itself releases a kind of power. I think that leads to a more complete state.

👦🏻 Koji

I know there was a six-month gap during your product manager years when you went to study animation?

👩🏻 Naonao

Yes. After leaving Tencent, I loved content creation and especially loved animation, so I wanted to learn it and see if I could work in that field.

👦🏻 Koji

Where did you study at that time?

👩🏻 Naonao

I enrolled in training courses. First an illustration class for character design, for several months. Then I learned 3D animation software like Maya. I found it extremely difficult because it was so hard to use.

Having been a product manager, I naturally resist poorly designed products. So I felt maybe I wasn't suited for this industry.

Plus, the animation industry as a whole — the pay, or rather, much of the time you're truly working purely for love. I didn't have the courage then. I felt I couldn't break into animation.

👦🏻 Koji

So after that you didn't pursue animation, and started your previous extreme sports-related venture?

👩🏻 Naonao

Yes. During that period I also met many people in the animation industry whom I deeply admired.

I remember visiting a modeling and VFX company. The whole place was very dim, no air conditioning. We chatted with a modeling director who'd been working five years for 10,000 RMB a month. But in his eyes, I saw this genuine light as he described his work. I was deeply impressed.

🟢 The Agent Moment for AI Video: Why Now?

👦🏻 Koji

Today we still want to focus most of our time on OiiOii. Roughly when did you first get the idea for this product?

👩🏻 Naonao

Making animation — I've wanted to do this since college, but never had the chance. The idea for OiiOii, I think, came while I was at ByteDance. It was 2022, and DALL-E 2 had just come out. I thought the technology was amazing and immediately connected it to animation.

The seed was planted then, but I didn't know how to approach it. Later I went to Bilibili to understand the animation industry, and my previous project "Lipu" — all of that was testing entry points.

First, understanding the industry. Second, testing where to切入. Only when I started my own company did I find that entry point.

👦🏻 Koji

At what point did you see a concrete entry point?

👩🏻 Naonao

Probably the first half of this year.

It was half passive, half active. Passive because at that point, it was becoming very difficult to continue pushing "Lipu" in the previous environment — I had to go out on my own. Active because we saw multimodal models starting to get intensely competitive, that momentum very similar to when language models exploded. It was good timing.

👦🏻 Koji

In building OiiOii, were you inspired by any AI products?

👩🏻 Naonao

Yes. Agent became hot starting this year. When we were thinking about how to approach animation, we realized Agent is a perfect form.

First, it can call various models.

Second, animation production itself is flow-based — its pipeline calls on various professional roles collaborating to complete a finished piece. Very suitable for Agent.

Third, Agent's interaction differs from traditional GUI. It gives creators much higher freedom. Traditional creative tools keep stacking features infinitely, but Agent doesn't.

Overall, Agent is the most matching form, so we used it for OiiOii.

👦🏻 Koji

Were there any model capability unlocks that particularly helped OiiOii?

👩🏻 Naonao

We planned in July, started building in August. The first version followed the same path as many Agents today.

If you're interested in AI video, you'll know there are two generation modes, or two "schools":

One is the "multi-reference school," using multiple images as references.

The second is "first-and-last-frame." Many domestic video models are especially good at first-and-last-frame.

We also used first-and-last-frame before, because stability within a single shot was better.

👦🏻 Koji

So if I have 8 shots total, I'd need 9 images connected end-to-end, then generate 8 short clips via first-and-last-frame, and stitch them together?

👩🏻 Naonao

Right. At that time we had an innovation: because each storyboard expresses different content, the suitable model should differ too.

For example, for action scenes, use Model A; for emotional character moments, maybe Model B works better.

We designed a Task Agent that knows each model's strengths and weaknesses, automatically matching the best model for that shot.

We built this version for a month, results were already quite good. We were adding TTS, background music, and sound effects when Sora 2 came out. I saw animation made with Sora 2 on TikTok and was shocked — completely indistinguishable from AI, just complete animation.

I thought this was amazing. We immediately decided to integrate Sora and see the results.

After integration, the results were excellent. The first video our R&D made was a small crab and small gorilla playing basketball. Very complete.

My first reaction was: wow, could we actually pull this off?! Very exciting.

👦🏻 Koji

That crab and gorilla basketball video — straight from one prompt?

👩🏻 Naonao

Straight from one prompt.

👦🏻 Koji

So no more of the previous first-and-last-frame tuning process?

👩🏻 Naonao

Actually our previous first-and-last-frame was also one-prompt output, just completely different generation method. Sora's expressiveness — one great thing is shot transitions. Cuts between shots are very natural, with lots of montage techniques.

But its weakness: it's reference-image plus text-to-video, so text weight is very heavy, leading to less stability in consistency and other aspects.

👦🏻 Koji

It has advantages in shot transitions, but because it's multi-image reference plus text-to-video, it's very text-dependent and consistency isn't ideal?

👩🏻 Naonao

Right, because it only "references" the images users provide, doesn't continue generating from that exact image. So its stability comes from the model's understanding. First-and-last-frame gives more stable results because it literally extrapolates from that first frame.

👦🏻 Koji

So after Sora 2 came out, what changes did you make to your product?

👩🏻 Naonao

We originally had two pipelines running simultaneously, then made a decision: stop the old pipeline, switch everything to Sora.

Because despite some issues, the results were excellent. It must have been trained on lots of films — when you generate or do secondary creation, you can see very heavy film influences, even tell which specific film. I found that impressive.

🟢 Why Aren't We Afraid of Sora?

👦🏻 Koji

If we get to Sora 4, Sora 5 in the future, could end-to-end Sora eat products like OiiOii?

👩🏻 Naonao

No, I actually think this is very friendly to Agent-type products.

My feeling from building Agent: video models are unlikely to unify — each has its own characteristics. Because data annotation standards, data quality differ, what they output will naturally have different traits.

Each model is like a large supermarket or market — they're your "ingredients." Building an Agent is like opening a restaurant using these ingredients.

For example, what we're building now is a Sichuan restaurant. Our target customers love Sichuan food, love spicy food, so I need to go to various markets to pick ingredients best suited for these people.

👦🏻 Koji

Sora 2 market.

👩🏻 Naonao

But the main work here is whether your own chef is good enough. Seasoning, controlling heat — about 60-70% of our work is tuning these very subtle things you can't see in the product.

Between large models and Agent, it's like supermarkets and restaurants. Users can buy ingredients from supermarkets and cook themselves, or go to restaurants for ready-made meals. This also means there can be many video Agents in the market. Video content is incredibly diverse, and each type has different production methods. So there can be Sichuan restaurants, Cantonese restaurants, Hunan restaurants, hot pot places too.

This will be a thriving state where everyone builds up the "food street" together — competitive, but more importantly, collectively prosperous.

👦🏻 Koji

That's a vivid metaphor — large models are supermarkets, OiiOii is a restaurant. I can buy groceries and cook, or dine out. But I'm curious: since Sora 2 can already end-to-end generate anime with shot transitions from one sentence, what's different about OiiOii's cooking?

👩🏻 Naonao

Take MVs as an example. Feed Sora 2 the same prompt and images, and it won't reach our level. Because we'll expand your single sentence — say, "have them do a K-pop dance to make an MV" — into a fuller form.

We found many well-made 2D MVs on the market, had models learn from them in our own way, built a knowledge base. When users input prompts, we call this knowledge base model to output polished prompts in its style, combined with shot language and sound effects, one-click output. The bulk of work is how to build this knowledge base.

Another example: story. There are actually some disconnects between our shots now. Many people say the continuity isn't good, but this is our choice. Because if you make every shot perfectly continuous, the whole story becomes very flat, no sudden turns, weak dramatic arc. You need to constantly tune in between to find the right state.

👦🏻 Koji

This is your "chef's" way of using ingredients?

👩🏻 Naonao

Right, it's about ingredient selection.

👦🏻 Koji

Any other similar examples?

👩🏻 Naonao

Yes, we actually have a recent example of "over-seasoning."

Some users had complained about inconsistent scenes, because Sora 2 is a text-weighted model. When we fed it data, we didn't include scene images — only scene descriptions in the prompts — which caused scenes to vary between shots.

To fix scene consistency, we started feeding it scene images. But we went too far. It became too consistent, making the whole frame feel rigid. Text-to-video models today need more imagination. If you feed them fixed images, their imagination gets constrained.

What happens then? The scene shouldn't steal attention from the characters. People should exist within the scene, not be pasted onto it. But with scene images, while consistency improved, the integration between characters and background felt off — like paper cutouts, something subtly wrong.

It's like when a customer says, "Why isn't your Sichuan restaurant spicy?" and you dump in chili peppers, only to hear, "Too spicy now." That kind of thing.

👦🏻 Koji

But if Sora 2 gets much stronger in the future, solving its consistency issues while also learning vertical-domain video knowledge — say, it builds its own MV knowledge base! Would you worry then?

👩🏻 Naonao

Not worried. Still using the supermarket analogy: Freshippo and Wumart sell groceries, but they've also started making their own prepared foods.

But prepared foods and restaurants are different. Restaurants have their focused clientele, with deep expertise in specific directions. That, I think, is inexhaustible.

🟢 Beyond "Video Cursor": Where's the Incremental Market?

👦🏻 Koji

So who is OiiOii, this "Sichuan restaurant," targeting right now?

👩🏻 Naonao

Our defined target users and actual users overlap in some ways, but there have also been surprises.

Before this startup, we built a UGC product called "Lipu," and considered entering the "motion comic" space, but quickly realized the timing wasn't right.

First, motion comics depend heavily on scriptwriting — not our team's strength.

Second, motion comics are a traffic-buying business.

Third, motion comics optimize efficiency within an established workflow and rely heavily on manpower, which we didn't want.

These three factors made it unsuitable for our current team.

Why not do another UGC product like Lipu? I think UGC isn't ripe yet. UGC content has short consumption value, insufficient information density — in today's information-saturated era, it can't sustain a large app. And UGC requires massive user volume, with high investment and inference costs.

So we're not doing UGC or motion comics for now.

We found our target audience in the middle ground: self-media creators — individuals or small studios whose content suits animated presentation.

This breaks into three categories:

First, people already making animation. Many work around a single IP, creating continuously. Our features let them build an IP character and produce content repeatedly and efficiently. We found that a small two-to-three-person studio currently updates about once a week. With our tool, theoretically they could update 10 episodes a day; even being selective, one or two episodes daily is completely doable — a massive efficiency gain.

Second, ACG creators making MV-style content.

Third, self-media creators outside animation — many history or science explainers whose content is perfect for animated presentation, except animation used to cost too much. These people have motivation to try animation, are willing to experiment with new tools, and already have traffic advantages on other platforms. They're not a niche group.

This is the audience we want to reach, and in our one-month beta, we did reach them.

👦🏻 Koji

Besides these expected users, who showed up unexpectedly?

👩🏻 Naonao

Right, some surprising users emerged too, also in three categories:

First, motion comic teams. We assumed our product couldn't meet their needs since we lack features like script upload built for motion comics. But their feedback was: they don't need such strong final-output capability. As long as the storyboards match, downloading all storyboards to edit themselves is enough. They found storyboard generation extremely efficient. These users are in our future plans, but it seems we can already serve some of their needs.

Second, people who've never made videos, don't like appearing on camera, but want to use animation to express their ideas and state of mind. Because animation is such an easy medium for expressing inner worlds.

Third, people creating for social relationships — not for mass audiences. Parents making things for children, students for teachers, couples for each other, people for their pets, adults for their parents.

👦🏻 Koji

Going forward, facing so many different user types, will you subtract and focus, or try to serve everyone?

👩🏻 Naonao

Going forward, we'll think a bit like how Douyin approached vertical categories. If we want to do science explainers, we'll find science creators, study all animated science videos, analyze their structures, turn that knowledge into our knowledge base, and serve them better.

Right now we're just at the "usable" stage, not yet "well-served." We see these promising signs, then prioritize and tackle them one by one.

👦🏻 Koji

So the ultimate goal is still to serve all these people well?

👩🏻 Naonao

Yes.

👦🏻 Koji

You mentioned Maya is too complex, and today many users find CapCut complex too. Do you worry OiiOii might one day become bloated?

👩🏻 Naonao

This is exactly why I think Agent is the perfect vehicle. Past creative tools, from Photoshop to Figma, from Premiere to CapCut, have all gone from simple to bloated, then been replaced by something simpler. Their core never changed — they just kept stacking features.

But Agent is different. It may accumulate capabilities, but won't get too bloated, because much of its power stays behind the scenes rather than showing through the GUI.

The charm of Agent is that it's a co-creation process between user and product. Users with exploratory instincts can discover uses we never imagined.

👦🏻 Koji

Users can find uses you didn't know about?

👩🏻 Naonao

Yes. Many of our current shortcomings are being solved by users figuring out workarounds. That's the charm of Agent products — it's not a dead thing. Users can keep exploring its boundaries, making it richer.

👦🏻 Koji

But Agent flexibility also brings unpredictability. Traditional editing software workflows are controllable, yet users often want both control and flexible efficiency. How do you balance these seemingly contradictory needs?

👩🏻 Naonao

That's a good question.

Mainstream Agents fall into two camps: one like language models with extremely high freedom, the other like n8n focused on workflow stability.

We want both pipeline stability and the ability for users to freely converse with the Agent at every step. This demands extremely high architectural design.

In just two months, we've iterated our architecture four times.

👦🏻 Koji

Four rewrites in two months?!

👩🏻 Naonao

Yes. Version one used System Prompts to define various Agents, letting the model judge workflow itself — but it was completely uncontrollable, the model wouldn't behave, too much freedom. Version two was strict workflow, but completely lost the freedom to modify.

Version three added a "signal" mechanism on top of workflow. The Agent could know when to "jump out" of workflow to accept modifications, then "jump back" via signal. Version four strengthened its freedom to jump to other mid-process steps on top of version three. We're still tuning between versions three and four.

This is an architecture prone to problems; we're still improving stability. Many issues users encounter now, like workflows getting stuck, are likely the Agent jumping out and forgetting to return, or never jumping out and just proceeding on its own.

👦🏻 Koji

You previously led CapCut, and now you're building OiiOii. Do you think video Agents will replace or eat into CapCut's market, or create an incremental one?

👩🏻 Naonao

I think it's incremental. CapCut can be divided into tool and ecosystem. As a tool, its track-based editing is replaceable. But as an ecosystem, it's grafted onto the entire Douyin, serving specific content formats — that's unique to it.

OiiOii is similar. Making MVs, making explainers — it's like building vertical knowledge bases, delivering specific content types. Hard to one-click generate with other tools. So it has uniqueness; it's not a pure editing process. Of course, this generated content can then be imported into CapCut for secondary editing, no problem.

👦🏻 Koji

So users generate content in OiiOii, then return to CapCut for post-production?

👩🏻 Naonao

Basically yes. We currently focus more on the front-end generation, because editing tools are already mature. Doing them well is very hard; spending massive energy there isn't worth it.

👦🏻 Koji

Today many entrepreneurs want to build "Cursor for CapCut" or "Cursor for video editing." What do you think?

👩🏻 Naonao

Before Sora, video generation mainly worked well within single shots; cutting between shots was weak, so editing mattered a lot — transitions, effects. I thought these editing capabilities would be hard for AI to replace. But after Sora 2 came out, we tried it and found these complex editing capabilities can absolutely be replaced by models.

But looking from the other direction, some light editing capabilities actually become irreplaceable. Like trimming heads and tails, or uniformly adding TTS (text-to-speech) at the end, and so on.

👦🏻 Koji

Makes sense when you think about it.

👩🏻 Naonao

Yes. Because when models learn from video, they already internalize many editing techniques, so they'll likely handle much of the heavy editing work.

👦🏻 Koji

Whereas micro-adjustments like deleting a 0.1-second shot — there's absolutely no need to operate AI through natural language for that.

👩🏻 Naonao

Right, completely unnecessary. You just drag the timeline, done.

👦🏻 Koji

What other examples are there?

👩🏻 Naonao

Just simple editing functions that don't need more complex alternatives — they're already simple enough.

It's actually the heavy, complex editing capabilities — adding transitions, doing effects — that are more easily replaced by models.

👦🏻 Koji

So it's possible people will still use CapCut alongside OiiOii — it's just that editing a video used to take three hours, and now it takes 30 minutes.

👩🏻 Naonao

Right, exactly. The optimal solution is combining both for efficiency.

👦🏻 Koji

OiiOii has another interesting design — you mentioned multiple Agents, each with their own name. I also noticed during the workflow that, for example, the script Agent might summon the character design Agent, like inviting everyone into a group chat. Could you elaborate on this design?

👩🏻 Naonao

From the start, I wanted to create a "team serving the director" feeling, since the product manager and director roles are quite similar. So each Agent needed a sense of character. The "group chat" is really just a fun visualization of the workflow process.

But this introduces a challenge with context memory across multi-Agent interactions — it's a dual-layer relationship. We only see these seven Agents, but underneath there may be other invisible Agent assistants doing the work. So architecturally, this is relatively difficult.

👦🏻 Koji

OiiOii's video quality is excellent. Beyond what you've mentioned, are there any other "secret ingredients"?

👩🏻 Naonao

For example, we want ordinary users to be able to make videos that "have a feeling." What's "feeling"? Sadness, joy, healing... these are emotional words. Ordinary people don't understand cinematic language, so how do we help them express something like "loneliness"?

We package extensive film and television expertise. For "loneliness," we might associate it with long corridors, gray-white color palettes, then combine these elements in various arrangements integrated into the generation process. So while the user only inputs an emotional word, the resulting video is backed by substantial professional knowledge — that's what gives it the "flavor."

👦🏻 Koji

It reminds me of Zhu Ziqing's The Back View — the way he expressed his father's desolation and care for him through that single image of a father carrying tangerines beside the railway tracks.

👩🏻 Naonao

Right. There are actually rational elements used to express something emotional — that's the part we need to get right.

👦🏻 Koji

Though sometimes artists also have those divine strokes of inspiration. Like Zhu Ziqing's scene — it's different from how anyone else would describe it, yet so delicate and impactful.

👩🏻 Naonao

That depends on the creator. We can learn from many great directors' styles — you input a director's name, and you get that flavor. But our current freedom isn't enough yet. There are still some advanced creators who are incredibly skilled — when they use our tools, I can't believe what they make.

🟢 WeChat's Right Brain, ByteDance's Left Brain: Two Cultivations for Product Managers

👦🏻 Koji

You're one of the few people who has worked in product at both the WeChat and ByteDance ecosystems. At these two top product factories in China, how did the experience of being a product manager differ?

👩🏻 Naonao

At WeChat, I was actually working on QQ Mail very early on, still at a junior stage as a product manager. That experience's biggest impact was cultivating very deep product values.

I saw how someone like "God Long" (Allen Zhang) could graft such profound thinking about human nature onto a product — that was incredibly powerful skill.

👦🏻 Koji

An example? How did his thinking about human nature translate to something as seemingly neutral and personality-less as QQ Mail?

👩🏻 Naonao

QQ Mail is a bit hard to use as an example. I remember when working on WeChat, every version update had a "preview." Once Xiaolong used the phrase "Everything I say is wrong," paired with a Michael Jackson image. It looks very simple, but you know there was extremely deep thinking behind it, and the final slogan that emerged was very powerful.

A product's sense of power often reveals itself in these subtle details — it really moved me.

👦🏻 Koji

That really was his divine inspiration. I remember when one WeChat version launched, it opened with the "Jump Jump" mini-game — completely unexpected that WeChat would make a box-jumping game, yet it was so addictive, everyone in the country couldn't stop playing.

👩🏻 Naonao

Right, there were many stories like that back then.

👦🏻 Koji

Just now you compared WeChat and Douyin's product systems. You mentioned feeling a strong "product values" sense in the WeChat ecosystem?

👩🏻 Naonao

At that time we went through user feedback every day — hundreds or thousands of pieces, and every single one had to be replied to. This had enormous impact on me. We had to identify from the feedback which were real needs and which were false — the loudest, most aggressive voices weren't necessarily the real ones.

👦🏻 Koji

Allen Zhang mentioned in his classic 8-hour product lecture how to identify true and false needs: the method is to "soak" in user feedback, see what users are actually saying, rather than sitting in conference rooms battling it out.

👩🏻 Naonao

Right. When we think too highly of ourselves, we become blind. So you have to "ground" the product in users' real scenarios:

First, see what they feedback; second, see their actual behavior.

The third point is very important: User experience is a "trained intuition." It's like a large model — through long-term interaction with user feedback, constantly reinforced, eventually forming an intuition.

👦🏻 Koji

What about at ByteDance?

👩🏻 Naonao

ByteDance is very data-focused. I was quite uncomfortable with it at first. Because good data doesn't mean good experience. For example, making a button bigger will definitely improve metrics, but that doesn't mean it's the optimal user experience.

I was somewhat resistant to this purely data-driven system early on.

👦🏻 Koji

Were you immediately the product lead for CapCut when you joined ByteDance?

👩🏻 Naonao

No, I started with effects, then CapCut, then worked at Douyin for a while. CapCut was actually great because it was an organization with stronger product experience sensibility.

It was really while working on effects that I gradually experienced the power of data and understood what "strategy product" means.

👦🏻 Koji

What's a strategy product?

👩🏻 Naonao

For example, we need to judge the highest probability of a user using a feature through various behaviors.

Because effects follow trends — to create a trend, you first need to find the people most likely to follow it. So who is most likely to follow trends? Someone who frequently opens effects, or frequently watches effect-based videos, or even just once clicked that spinner below, or clicked save — using various behaviors to calculate a probability of whether they're more likely to use effects.

When you understand the principles behind this, you discover that product sense and data can be perfectly combined.

It's not cold numbers, but a logical deduction process backed by real user behavior. You need to know how to read data, and how to use data.

👦🏻 Koji

Pretty interesting.

👩🏻 Naonao

Right. One is very right-brained, one very left-brained. This is very helpful for me now making AI products. Because AI products don't have massive data volumes for you to make judgments — you need that product sense cultivated in the WeChat system; but at the same time, handling model coordination and strategy requires that rational thinking trained in the ByteDance system. It's like animation — a perfect fusion of art and technology.

👦🏻 Koji

That's indeed an interesting metaphor — making products seems like right brain at ByteDance and left brain at WeChat, yet both made China's most powerful products. So what do they have in common?

👩🏻 Naonao

Their commonality is taking their own strengths to the extreme.

ByteDance focuses on data science — its recommendation engine and growth system take the advantages of "efficiency" and "data" to the extreme.

WeChat at that time had very deep understanding of users and human nature. It didn't洞察 surface-level needs, but rather what deep human nature浮上来后 should be承载 through what kind of experience. It took "experience" and "perceptiveness" to the extreme.

👦🏻 Koji Like how back then no one expected WeChat's slogan would be "WeChat, a way of life."

👩🏻 Naonao

At WeChat, it felt like Allen Zhang was making decisions. But he wasn't closed-minded or stubborn — he openly listened to suggestions, he just usually thought deeper than others. The final decision was a convergence based on hearing everyone's opinions.

At ByteDance, people believed more in "scientific decision-making," believed in the organization, or rather believed in the answers given by this algorithmic engine.

👦🏻 Koji Though the methods differ, the core is quite interesting: Everyone realized where their advantage lay, then amplified that advantage to the extreme.

👩🏻 Naonao Right. It's like entrepreneurship — there's no single guaranteed path to success, every founder has their own personality.

The key is knowing your strengths, then taking that direction to the extreme — that may be how you create something with rhythm, something that can exist long-term.

👦🏻 Koji

So what do you think OiiOii's advantage is?

👩🏻 Naonao

First, from a third-party perspective, founder and team DNA is extremely important. You must deeply understand and truly believe in both animation and technology.

In the business world, pure passion doesn't necessarily lead to commercial success. But animation is an exception (like Disney, Pixar, Light Chaser's Gary Wang, Jiaozi) — it's one of the few industries that rewards "purity" and "passion." This suits me very well.

👦🏻 Koji To summarize, where do your advantages as an entrepreneur lie?

👩🏻 Naonao

First, making animation has been a very, very long obsession for me — I didn't just decide to do this when AI arrived. It's something I've always wanted to do; AI just gave me an opportunity to do it better.

Second, it's the match of capabilities and team. I've been at large companies, I've founded my own startup — to some extent, both I and my team are prepared for this. Having both conviction and capability — then just go do it.

👦🏻 Koji

What do you think is the most important thing for being a good product manager?

👩🏻 Naonao

First is empathy. You need to rapidly switch roles — even with deep industry experience, you must be able to instantly become a novice again. This requires you to frequently step back, observing yourself and users like an observer.

Second is 50% confidence and 50% self-reflection. Confidence without self-reflection easily becomes arrogance. Self-reflection isn't for self-abasement, but to more calmly see shortcomings, thereby strengthening confidence. This reduces your ego — and thus reduces the fatal blind spots of product managers.

Third is technical sensitivity. You don't need to code, but must have good understanding of technology, knowing what technology can achieve what. This is basic competence.

👦🏻 Koji

How do you maintain technical sensitivity?

👩🏻 Naonao

I don't seem to deliberately maintain it — it's like intuition, something that gradually accumulates through interest.

For example, I was visually and auditorily sensitive as a child — that's the perceptual side. Second, I really liked physics, so I'm very sensitive to rules. These things all lead me to be sensitive to multimodal technology, which happens to be audiovisual language technology with patterns — it's precisely my strongest area, so I became sensitive.

👦🏻 Koji

It feels like you're in a very fortunate position now — the things you were interested in as a child, the things you wanted to do when you were young, you can finally do them.

👩🏻 Naonao

Yes, with this startup I have a very strong feeling of luck. The detours I seemed to take in the past were actually preparing me, step by step, for this moment.

👦🏻 Koji

It's like that Steve Jobs quote: connecting the dots backwards.

👩🏻 Naonao

Right. I wasn't deliberately arranging things this way before. But you realize that this time, the moment has arrived — my abilities, my passions, my sensitivities, they all align. Everything is just right for doing this. It's wonderful, just very lucky.

👦🏻 Koji

Where did the name OiiOii come from? It's quite distinctive.

👩🏻 Naonao

"Oii" sounds like a greeting in anime culture — very friendly and cute. It was suggested by one of our engineers. We went with OiiOii because two syllables sound even cuter, and if you look at it written out, it resembles two little snails. We hope to be like snails — steadily, slowly climbing, but also adorable.

👦🏻 Koji

That's fascinating. My initial understanding of "oii" was that it rolls off the tongue easily, a bit like an anime greeting, very energetic.

But I didn't realize there's not just sound behind it, but also a layer of pictographic meaning.

👩🏻 Naonao

Right. That's why the spelling is "one o, two i's" — it looks more like a snail that way.

🟢 Conflict as a Force for Doing

👦🏻 Koji

You once shared on Moments: "Manifesting invisible willpower into a product is the secret technique — few have mastered it, and I am far from pure enough." What story, product, or person moved you to write that?

👩🏻 Naonao

I was preparing to leave "Lipu" to build something new. Giving up Lipu was hard for me — it was my starting point for making animation. But in the process of building it, there were many things beyond my control, too many "impurities" mixed in. I felt that with this startup, even though I hadn't found a great entry point yet, my determination to do animation was certain. I could try it in a new product, using only the force that I myself directed, to do something more pure.

👦🏻 Koji

Hoping this willpower comes more from your pure self.

👩🏻 Naonao

Or from the team I've chosen.

👦🏻 Koji

In the same post, you also wrote: "You have to allow conflict, even create it, because conflict is the force for doing things, the way to filter who is truly doing things." Can you share a story where conflict became a force?

👩🏻 Naonao

I love competitive sports. Playing basketball, for instance — you find that opponents often bring out your potential.

Healthy competition is a process of mutual appreciation and mutual inspiration.

But you rarely encounter this state at work, because usually when you're in conflict you're in attack mode. Yet we have had teammates we came to know through fighting.

For example, at ByteDance, multiple teams were competing for the "effects" business. We ended up winning it, which created conflict with other teams. But later we discovered that since everyone was "putting the work first," after fighting we actually admired each other — opponents became teammates.

During my first startup, the first person I fired was a good friend, because his energy couldn't keep up with the team. When I told him, he cursed me out brutally. After leaving, he even poached our people to do the same thing. But a year later, he added me back on WeChat and said: "I finally understand you." Because once he was in that role himself, he experienced what I had been going through.

These experiences gave me positive feedback: conflict at the time may cause misunderstanding, but as long as your intention isn't to harm others, time will eventually prove everything. Sometimes, conflict actually inspires a greater attitude toward doing things in each other.

👦🏻 Koji

Why do people call you "Lord Nao"? You seem very peaceful today — quite a contrast with the character "nao" (noisy/restless). Were you noisy in the past and only recently became peaceful?

👩🏻 Naonao

I was very noisy in the past — a very extroverted kid. Maybe more than ten years ago, during my last startup, I went through some ups and downs, and some things happened in my family. This made me more at peace. In this process, I tasted the sweetness of peace and discovered it's actually a greater power.

👦🏻 Koji

What kind of sweetness, for example?

👩🏻 Naonao

In middle school I formed a band, played rock, later got into extreme sports. My inner self was very rebellious, with a defiant streak. You'll find that kind of power is outward — shouting, resisting, trying to seek a high degree of freedom from the outside world.

But no matter how you search externally, that state is never satisfied. You think that's freedom, but it's actually a huge cage you've built for yourself, constantly consuming you.

When you have the chance to examine your inner self, you find true freedom is here. This energy isn't forceful or explosive, but continuous and flowing — it's actually more expansive. People think "noisy" and "quiet" are very contrasting, but to me they're two sides of the same coin, both seeking freedom — at first I sought it through shouting, and finally discovered true freedom is within.

👦🏻 Koji

I really hope we can talk more about this transformation from "noisy" to "quiet," and how to gain more peacefulness to generate power.

But back to OiiOii — you said your favorite entrepreneur is the founder of Pixar. He started with technical tools, then pivoted to animation and achieved enormous success. Why didn't you, like your idol, go directly into making animated films?

👩🏻 Naonao

I admire him not because he made animated films, but because he found what he was good at and pushed it to the extreme in animation.

We instinctively think making animation requires being able to draw or tell stories, but he was good at neither.

What he loved was computers and physics — same as me. He found his way through gradual exploration. For example, using graphics and mathematical fractal techniques, he created the world's first computer-simulated hand. He didn't use traditional methods, but approached his passion through technology.

👦🏻 Koji

Right, that one-minute short about the hand caused a sensation when it was released.

👩🏻 Naonao

Right, and it opened the door to creating virtual images with computers.

When he was doing special effects at Industrial Light & Magic, he realized that even without being able to draw, using his expertise in computers and physics, he could still participate in animation production — which eventually led to the world's first 3D animated film.

I can draw, but not well enough; I want to write stories, but can't write earth-shattering ones; I'm not from a technical background either.

But I have my strengths — I'm sensitive to technology and know how to translate it into products. My satisfaction comes from people using this product and having their creativity sparked.

So in building OiiOii, it's using what I'm best at to engage with the animation industry. That's where the inspiration lies for me.

👦🏻 Koji

What do you hope OiiOii becomes in the future? Remain a tool, become some kind of content platform, even become the next Pixar?

👩🏻 Naonao

I hope it has more possibilities beyond being a tool, but I don't want to nail down the future. I hope it can at least enable anyone who wants to make animation to create their own animated film with it. Once we achieve that, we can explore other possibilities.

Those possibilities are fuzzy concepts in my mind, but I don't want to define them — once defined, they collapse.

🟢 Predictions for 2026

👦🏻 Koji

What changes do you think are highly likely in AI video by 2026?

👩🏻 Naonao

The trends of the past year or two are obvious — quality will keep improving, editability will keep improving. Going forward, real-time capabilities may strengthen, bringing some interactive elements. But I don't think that's a huge revolution, because ultimately it depends on whether audiences buy in.

For example, the higher the editing freedom, the more professional the audience becomes, and the user base actually shrinks.

Interactivity is an active behavior, but we've already been conditioned by short videos into passive information reception. The moment you ask users to take one extra step, the crowd becomes very narrow.

So my other judgment is: I don't set limits.

Look at Sora 2 — it just achieved more natural shot transitions. That doesn't sound like a huge change, but the effect is excellent. Even without major innovation, without interactivity or editability, just making the consumption experience slightly more natural, it works very well. So perhaps many predicted revolutions won't bring much change to actual consumers, while small tweaks on existing media may reach larger audiences.

👦🏻 Koji

You mentioned OiiOii matches different models for different shots — Vidu, Keling AI, what are they each good at?

👩🏻 Naonao

For example, in animation there are fight scenes — we'll call "Hailuo AI," which is better at that.

If you need delicate facial expressions and emotions, Seede might perform better.

If you want that CG blockbuster feel, "Keling AI" would be more suitable.

We automatically assign based on shot descriptions.

👦🏻 Koji

Do you think these video model teams are deliberately pursuing their own differentiation?

👩🏻 Naonao

I think there's innate and acquired differentiation.

Innate — for example, initial data and annotation standards differ, so the output is completely different. Because there are many steps in between, and humans are most sensitive to visuals, you immediately feel the difference in output.

Acquired — this may relate to each company's strategy. If you want to make cinematic blockbusters, you'll train more on that type of material, so naturally you'll be strong in that area.

But video models are hard to unify. Even with identical models, different prompting techniques, different inputs, produce completely different results.

Take our previous experience tuning Sora: we used the exact same Sora model, but initially the scene consistency wasn't strong. The reason was, in our input we only gave a single small element, not a scene image. Although the model foundation didn't change, this tiny change at the input end led to completely different outputs. Not to mention that each model differs in training data dimensions from the very beginning — this inevitably creates huge differences between models.

👦🏻 Koji

Can users choose their own models in OiiOii?

👩🏻 Naonao

Not for now, we might add it later.

👦🏻 Koji

How do you personally judge the competitive landscape among these video model companies in 2026 and beyond?

👩🏻 Naonao

I think everyone will move in two directions: strengthening what they're already very good at, and filling gaps in what they're not good at.

Model companies will still pursue greater generality. Plus the two major directions I mentioned: real-time capabilities and editability — these may significantly strengthen.

👦🏻 Koji

Alright, last question. If you had $3 million for angel investment, split into three portions to invest in three people — whether they haven't started a company yet, are currently working, or are already founded and want follow-on — who would you invest in?

👩🏻 Naonao

My first instinct was to think of many people, but my second thought was: I'd still want to invest in myself.

It's not because I think I'm anything special. From a third-party perspective, I'm simply the person I know best. Angel investing starts with finding someone you're truly bought in on — and I'm the person I'm most familiar with.

🟢 Ten "I am ___" Statements

👦🏻 Koji

Then let's end with a bonus round. Naonao, please introduce yourself using ten "I am ___" statements.

👩🏻 Naonao

That's hard. For me, this question is especially difficult because I'm actively trying to reduce my accumulation of "I am ___" identities. I've studied some Buddhist teachings and understand that stacking or reinforcing "I am ___" brings many problems. For example, if you label yourself "I am a professor," you have to maintain that image — it takes so much effort. So in my current state, this is genuinely hard.

👦🏻 Koji

Very profound. So your first answer is "I am someone who is actively trying to minimize 'I am ___' statements."

👩🏻 Naonao

Yes.

👦🏻 Koji Then I'll exempt you today — you don't need to say many, just share what feels most important.

👩🏻 Naonao

My first impression is that one. And my second — if I must say another — I think I am a "vessel." This connects to OiiOii. I feel there's a "god of animation" that simply expresses things through me as a vessel. It wants better forms of expression in this era. It's not "me" expressing — something wants to express, and this vessel happens to suit it.

👦🏻 Koji

OiiOii itself doesn't express, but helps those who want to express do so better — it's a vessel?

👩🏻 Naonao

Yes, it's a vessel, and so am I. This inspiration came from an Ang Lee interview. He said there are plenty of people more talented than him; they didn't achieve what he did not from lack of talent, but because the "god of cinema" happened to choose him. He became a vessel through which film could express itself.

In some ways, I feel this too. I'm not trying to express something — I hope to become a channel for everyone else's expression.

👦🏻 Koji

Thank you, Naonao. I hope we can talk again before long, and see what new stories have grown inside this vessel that is OiiOii.

👩🏻 Naonao

I hope so too. Thank you.