A Conversation with PixVerse Co-founder Xie Xuzhang: When Video Foundation Models Enter the Final Stretch

The Era of Everyone-as-Director Is Almost Here

👦🏻 Interviewer: Koji

🥷 Editor: Crossing

🧑‍🎨 Layout: Zeoooo

On June 12, at the AltNext conference, we sat down with Jaden Xie, co-founder of PixVerse, for a conversation about video foundation models: the breakneck competition among Chinese companies, the wave of American players quietly exiting the field, and the roadmap from PixVerse R1 to the upcoming R2 — how the race has entered its final stretch, what comes after the Seedance era, and how ordinary people might finally claim the right to create video.

👦🏻 Koji

One theme here today is "In the future, everyone can be a director." As a founder on the front lines of video foundation models, Jaden — how many days until that future arrives? 500 days? 1,000? Or 30?

🧑🏻‍💻 Jaden Xie

We're actually very close already.

AISphere started working on video generation in 2023, when it was still a niche direction. Our product path has been somewhat unusual — we focused on ordinary people from the start, lowering the barrier to creation.

We now have over 100 million users globally, and research shows that 50% of them had never made a video before. This means anyone with a phone can become a video creator.

What we need now is model iteration and collective industry effort to make the public aware that AI is ready — ready to help everyone become an AI-native creator.

You can't measure this in days anymore. It's right in front of us. We just need to push it into everyone's hands.

👦🏻 Koji

You're remarkably optimistic. You believe that day is inevitable?

🧑🏻‍💻 Jaden Xie

Yes.

Signals

👦🏻 Koji

Along the way, what concrete signals or stories have struck you?

🧑🏻‍💻 Jaden Xie

Two moments hit me deeply.

The first came just months after our product launched. A retired cadre in his sixties or seventies sent us a long email. We were running a campaign recruiting AI creators and giving away free tokens, and he signed up.

He wrote that he'd harbored a director's dream his entire life, but the tools and editing barriers had always kept it out of reach — until AI lowered the threshold.

That letter moved me. There are so many people in the world who watch video every day yet have never created content with their own hands. AI is genuinely giving everyone the chance to become a creator.

The second moment came from user research we conducted before Chinese New Year. The data showed that over 50% of our users had never made a video before — meaning they'd never created video anywhere outside of PixVerse.

👦🏻 Koji

So they hadn't used CapCut or any other tool to make video?

🧑🏻‍💻 Jaden Xie

Correct. Billions of people watch video every day worldwide, yet the overall submission rate on short-video platforms is under 10%. That means 90% of people can only consume content — they have no creation experience whatsoever.

Video has become the dominant information medium, but the right to create has never truly been democratized.

👦🏻 Koji

How did the retired cadre's work turn out?

🧑🏻‍💻 Jaden Xie

He made his first feature-length piece.

We gave him plenty of free tokens. For a complete beginner, the result was quite good.

Video is fundamentally a multimodal art — story, image, editing. I saw that he kept at it, using our tools and others, and he's still creating today.

👦🏻 Koji

Speaking of "everyone being a director," Jaden — as a co-founder, have you tried directing videos yourself?

🧑🏻‍💻 Jaden Xie

In our first year, I thought: I'm building a video AI company, but I've never actually made a video myself. So during Qingming Festival, I visited the Daqing Museum and took lots of photos of dinosaur fossils, then used image-to-video to "resurrect" them all.

I spent a full week learning CapCut — voiceover, transitions, the works — and cut together a video I was incredibly proud of. I posted it, and after one week it had exactly two likes. Almost no one watched.

That's when I realized: making video is still extraordinarily difficult.

👦🏻 Koji

You didn't use your company's distribution channels to promote it?

🧑🏻‍💻 Jaden Xie

No. So this made me realize: if the creative workflow is too complex, the barrier remains too high for ordinary people.

This shaped our core product thinking going forward. Users don't want complex editing — they want simpler, more direct expression.

So we defined a core product mechanic: users upload a 5-to-10-second real-world video, and the system extends the final frame through image-to-video, creating an extreme contrast between reality and the virtual.

One classic case: a user filmed their backyard dog walking for 5–10 seconds, and at the 10-second mark the dog suddenly put on a suit. That video got nearly a million likes on TikTok.

What ordinary people need is a simpler, more direct, more foolproof way to turn their lives into shareable, high-quality video content.

👦🏻 Koji

Rooted in life, but transcending it?

🧑🏻‍💻 Jaden Xie

Exactly. It's like helping people do a remix of their lives. You take the ordinary, elevate it through remix, and share it.

Global AI Video: Chinese Players Are Too Fierce, Too Good

👦🏻 Koji

After Seedance came out, what did you and your peers discuss most? What made you most anxious or excited?

🧑🏻‍💻 Jaden Xie

Seedance is an unavoidable benchmark for the entire industry — exceptionally strong, generating nearly $200 million in monthly revenue globally.

👦🏻 Koji

And that's with limited GPU supply.

🧑🏻‍💻 Jaden Xie

Right. And video generation is very different from code or LLMs. Last month I was talking with overseas peers, and there's a strange phenomenon in this industry: around the time Seedance emerged, the market expanded nearly 10x, yet the number of deeply engaged players is shrinking.

Many American peers aren't making significant progress. Including OpenAI and Google — I attended Google I/O, and their Omni model's video quality was actually below expectations.

Runway, one of the earliest video generation companies globally, has been getting quieter.

Though Luma is pushing its "world model," overall there are fewer people working on this overseas, even as the market space is expanding explosively.

👦🏻 Koji

What's causing this exodus from the video model battlefield?

🧑🏻‍💻 Jaden Xie

Chinese players are too fierce.

Whether it's ByteDance's Seedance, Kuaishou's Keling AI, us, or Alibaba — everyone's iteration speed is accelerating dramatically.

In talent density and technical reserves, China's accumulation in video runs deeper than overseas. Over the past 5–10 years, the vast majority of the world's most influential video products and applied technologies were built by Chinese teams. That underlying concentration is extremely high.

So our top priority lately has been self-iteration, polishing the product, and capturing our share of this rapidly exploding market.

The second topic is real-time video foundation models.

In January we released the first real-time world model, PixVerse R1, and we're preparing to release the second generation. We're incredibly excited: when humans can interact with content in real time, when "creation" and "generation" become one, what entirely new scenarios and products emerge? That's our core discussion.

👦🏻 Koji

One is maintaining innovation and rapid iteration in the face of Seedance's fierce competition; the other is exploring scenarios for the next generation of PixVerse R1?

🧑🏻‍💻 Jaden Xie

Yes.

👦🏻 Koji

With giants like Seedance and Keling AI dominating, how specifically does PixVerse break through in the next phase of competition?

🧑🏻‍💻 Jaden Xie

When we started, I drew a four-quadrant chart: domestic/global, B2B/B2C. Most peers entered either B2B or the pro-creator segment of B2C — tools for film-grade or advertising-grade creators.

Our differentiation was focusing earliest on serving "absolute beginners" with zero video experience, helping them make their very first video. That's a path that remains deeply non-consensus even today.

👦🏻 Koji

Why haven't peers followed this direction?

🧑🏻‍💻 Jaden Xie

Honestly, I don't know.

Sora had tried a consumer-facing app and later shut it down, which created an industry-wide perception that "C-end tools don't work." But in fact, our volume here is massive. By MAU and cumulative users, we're already one of the largest video generation platforms globally, helping tens of millions of ordinary users make videos every month.

The reason for this perception gap: language models follow a "model-as-product" logic — a chat box where base model capability determines the experience ceiling. Video generation logic is completely different.

Video has extremely diverse categories — short dramas, features, consumer play, enterprise use — each with completely different workflows and interaction logic.

Legacy players like Runway positioned early for film-grade professional content, which caused most companies entering the space to habitually chase 4K professional quality. As a startup, we iterate fast and are willing to try many directions. We built a new business line serving ordinary people, which let us find different entry points.

👦🏻 Koji

What are the key upgrades in the next generation of PixVerse R world model compared to R1?

🧑🏻‍💻 Jaden Xie

Dramatically faster response times. While R1 was already near real-time, it still had a few seconds of latency. The next generation will achieve sub-second response.

It's not just speed — it's a qualitative shift in interaction. As response time approaches zero, users can control the video generation process in real time through keyboard, mouse, or joystick — directing orientation, angle, and depth.

At the same time, we're evolving from 2D perspective toward 3D physical world models, and beginning to build an AI-native interactive game engine.

Real-Time Generation: Using AI to Kill Traditional Rendering

👦🏻 Koji

An AI-native interactive game engine? How should we understand its scenarios?

🧑🏻‍💻 Jaden Xie

We recently had deep discussions with a former Unity technical executive. In January Google released Genie 3, and we released R1, which sparked massive industry attention. We believe future visual content creation, including game visuals, may no longer require traditional game engines with their long bake and render cycles — instead, AI can generate directly in real time.

Games themselves are collections of visual content. Future video games will use "real-time generation" to replace "traditional frame rendering." This will open up an entirely new form somewhere between video and games — that's the focus of our next exploration phase.

👦🏻 Koji

R1 had extremely high compute load at launch and only did small-scale beta testing. Is it now fully open to the public?

🧑🏻‍💻 Jaden Xie

Yes, it's available to all users now.

👦🏻 Koji

Beyond speed, will the new generation be more efficient?

🧑🏻‍💻 Jaden Xie

Parameter count has increased significantly, so training and inference do consume more resources. But we have extremely strong engineering optimization capabilities. With R1, from launch to today, we've optimized algorithmic and compute efficiency by over 10x.

The next generation will follow a similar rhythm: first use beta testing to accumulate engineering optimizations, then scale to mass commercial deployment.

👦🏻 Koji

On the training data side, what additions have you made for the next generation?

🧑🏻‍💻 Jaden Xie

Training on 3D physical space and real-world physical commonsense data.

👦🏻 Koji

Is the data source mainly real-world capture, or synthetic data from game engines and simulators?

🧑🏻‍💻 Jaden Xie

Both. We've captured large amounts of real-world footage with panoramic cameras, while also generating rich synthetic data through high-precision game engines. The complementarity of synthetic and real data lets the model expand into richer use cases.

👦🏻 Koji

During R1's beta testing, did any user scenarios exceed your expectations and produce surprising work?

🧑🏻‍💻 Jaden Xie

During beta, nearly ten thousand people were interacting deeply every day. We'd designed many interesting official scenarios, but data showed close to 60% of users ignored the official scenarios entirely — they were all doing their own UGC.

👦🏻 Koji

This is what you mean by "creation as consumption"?

🧑🏻‍💻 Jaden Xie

Yes — creation itself is their consumption experience.

Some users would casually snap a photo during business trips, upload it, and use the large model to generate magical transformations of that scene in a parallel universe. The creation process is purely for self-satisfaction, not external distribution — yet the joy is immense.

Others would upload game screenshots and have the model extend outward, experimenting with what happens without game rule constraints. This has informed us that in productization, we need to further unleash users' UGC creativity.

👦🏻 Koji

Like people going to a pottery studio to make ceramics — the act itself is creation and consumption. In the future, with low-barrier tools, we might create a game for ourselves, and that process itself is the purest joy.

🧑🏻‍💻 Jaden Xie

Exactly.

👦🏻 Koji

From a底层 model and algorithm perspective, what are the different challenges and technical emphases between serving "ordinary users" versus "professional creators"?

🧑🏻‍💻 Jaden Xie

The most fundamental difference lies in device and tolerance for waiting time.

Ordinary users overwhelmingly consume and share video on mobile, so we must treat mobile as our main battlefield. But on mobile, most models' generation speeds can't meet instant experience demands — at the time, generating a 5-second video typically took 1–2 minutes. Users on mobile simply won't wait that long; they'll swipe away immediately.

So for our ordinary-user-side model R&D, the focus is on "extremely fast." We optimized an ultra-speed model for mobile, achieving "5 seconds to generate 5 seconds of video." Users upload a photo, pick a template, get results in 5 seconds, and share with one tap — massively shortening the feedback loop.

Because we achieved 5-second generation a year and a half ago, we dared to push further: if we can do 1 second, or even faster, can we achieve real-time video interaction and infinite generation? That's the driving force behind our exploration of the R1 real-time foundation model.

👦🏻 Koji

In your products, besides using PixVerse's own foundation model, do you integrate third-party models for specific scenarios?

🧑🏻‍💻 Jaden Xie

Our product line has four main segments: mobile app for C-end users, web端 for professional creators, enterprise platform, and real-time interaction model R1.

On mobile we only use our own models. But on the web端, we integrate third-party quality models. Professional creators' workflows are extremely complex; we bring in external quality tools to better match their needs.

👦🏻 Koji

Your massive global user base generates enormous data daily. What role does this data play in retraining your models?

🧑🏻‍💻 Jaden Xie

The role operates on direct and indirect levels.

Directly, massive real feedback data enables reinforcement learning. We've accumulated extensive multi-dimensional feedback signals including downloads, favorites, and likes, which lets the model precisely perceive which video qualities users truly value.

Indirectly, data serves as a "compass," guiding our product iteration and model pre-training direction. For example, when we discovered users on mobile have extreme demands for generation speed, we rapidly adjusted compute allocation and made speed our top priority in the next pre-training cycle and engineering optimization.

Three Years of What Changed and What Didn't

👦🏻 Koji

AISphere was founded in 2023. Standing here in 2026, looking back at three years of radical change, what underlying beliefs about video generation have remained constant, and what has been completely overturned?

🧑🏻‍💻 Jaden Xie

What hasn't changed is our original conviction. In early 2023, when the industry was just starting and video generation was an extremely non-consensus direction, we firmly believed AI would ultimately revolutionize content creation. That bottom-up confidence remains unchanged.

👦🏻 Koji

What specifically made this "non-consensus" at the time?

👑🏻‍💻 Jaden Xie

At that point, no team globally had yet produced a high-quality video generation foundation model.

LLMs only became credible after ChatGPT launched.

👦🏻 Koji

People "believed because they saw."

🧑🏻‍💻 Jaden Xie

Right. Our team chose to "see because we believed" — making a heavy bet while most hesitated. We believed AI would help produce new content creators, new content forms, new creator communities, new platforms.

👦🏻 Koji

So what changed over these three years?

🧑🏻‍💻 Jaden Xie

What changed is that we completely underestimated the industry's development speed.

Initially we predicted that from technological germination to the "everyone a creator" era would take at least 5 years or more. But the reality is, the past year or two the industry has been racing forward at an almost uncontrollable velocity.

From Sora, Keling AI, our models, to later Google and Seedance 2.0 — the entire industry is still accelerating. What was once "non-consensus" has in extremely short time gathered ever more confidence from giants and capital.

👦🏻 Koji

This compressed industry cycle — what pressure has it brought?

🧑🏻‍💻 Jaden Xie

It means you have to be brutally hard on yourself.

👦🏻 Koji

Under such intense competition, when do you think this race reaches its conclusion?

🧑🏻‍💻 Jaden Xie

I think we're roughly in the final stretch now.

The companies with the resources and capability to train next-generation top-tier video foundation models are already countable on one hand globally.

👦🏻 Koji

In this final stretch, what's the core winning factor?

🧑🏻‍💻 Jaden Xie

The most foundational, immovable factor is still high-quality models.

With base model quality above the threshold, then it's about competing on extreme engineering, productization, and commercialization capabilities — these determine how high the model reaches commercially.

End of 2026: Awaiting the Next Truly Breakout Real-Time Scenario

👦🏻 Koji

Final question. It's June 2026. If you predict the industry's development by December, what changes not yet consensus today will have become normal in six months?

🧑🏻‍💻 Jaden Xie

I believe by year-end, Seedance 3.0 may officially debut, and globally there should be one or two top-tier video foundation models fully on par with it.

👦🏻 Koji

Where are these potential peers most likely to emerge?

🧑🏻‍💻 Jaden Xie

Most likely all in China. We're going all out, hoping to be one of them.

Additionally, in the real-time video generation direction, I'm extremely hopeful that by year-end, the first truly breakout consumer scenario will emerge. Like a few years ago when an AI video "superhero transformation" effect went viral globally.

We hope in real-time interactive video, we can soon find the first globally resonant scenario.

👦🏻 Koji

Very much looking forward to that iconic moment, when everyone can become the director of their own story.

Thanks for sharing, Jaden, and thanks to everyone here.

🧑🏻‍💻 Jaden Xie

Thank you, Koji. Thank you, everyone.

Crossing is seeking independent writers to cover AI product and model reviews.

If you've written articles like: "Hands-on with PixVerse C1", "Hands-on with LibTV", please contact zeo0811@gmail.com. Your email should include: ① personal introduction, ② AI review articles you've written.

We offer competitive compensation. Looking forward to observing and documenting the AI era together 🎪