What Is the Next-Generation Real-Time Interactive Model? Vivix's Liu Yu on First-Principles Thinking | BlueRun Ventures Headline

**"Interactive content is an entirely new content category born in the AI era. If a model can only do audio and video generation, it's merely making the production of legacy content formats more efficient and cost-effective."** That's how Vivix, an AI real-time generation model company, sees the next evolution of content.

"Interactive content is an entirely new content category born in the AI era. If a model can only do audio and video generation, it's merely reducing costs and improving efficiency for production pipelines built around content formats from the previous era." This is how Vivix, an AI real-time generation model company, sees the next evolution of content.

BlueRun Ventures led Vivix's Pre-A round. From our perspective, what makes Vivix worth watching isn't just its ambition to lead the next generation of interactive content ecosystems, but that the team combines model R&D with infrastructure-building capabilities, iterating at extraordinary speed. Founder Yu Liu has deep experience leading large-scale multimodal model development, paired with the ability to build complete technical architectures around long-term goals.

Vivix aims to achieve "real-time interactive content," and is continuously building out its models, data pipelines, and infrastructure. In July this year, it released models W1 and A1 — W1 is a foundation model integrating understanding, interaction, and generation; A1 is a real-time interaction model centered on character interaction. Going forward, Vivix also plans to advance data scaling for W2, as well as model architecture and MoE scaling for W3.

In the in-depth conversation below, Liu shares his thinking on Vivix's journey so far: Why interactive content? Why shouldn't AI-era models merely optimize content production? He also shares his perspectives on model scaling, multimodal technical approaches, real-time interaction, entrepreneurship, and directions of technological evolution.

Here is our deep dive on Vivix — enjoy reading:


What Kind of AI Content Would People Rather Consume Than Douyin?

By Deng Yongyi

Edited by Zhang Yuxin

What exactly is Vivix doing? For more than a year, this has been the biggest mystery surrounding this enigmatic company in AI circles.

The impression Vivix gives the outside world goes something like this: star founder Yu Liu, aggressively hiring across every business line from algorithms to product to brand, company valuation soaring to $1.32 billion within a year, backed by a roster of illustrious investors.

But few can clearly explain what the company actually does. Some say it's a pure AI product company that doesn't build models, aiming to become "the Douyin of the next era" — which would explain the steady stream of interactive entertainment products like 7verse and TipTap. If you searched for Vivix during this period, the only visible technical direction on its website was "next-generation visual generation engine."

In the 19 months since Vivix's founding, competition in the multimodal space has already become as intense as in large language models (LLMs), with success stories like Seedance emerging.

"Are we that mysterious?" When I first met Vivix founder Yu Liu in July, he seemed somewhat oblivious to the various outside assessments. "We appear mysterious because when the model isn't very ready yet, there's not much point in PR."

Liu says that Vivix's previous products were all explorations to validate the model, and what he wants to do has never changed: next-generation real-time interactive models.

Much of the industry's attention on this company stems from some eye-catching labels: founder Yu Liu, born in 1995, a thoroughly bred SenseTime talent who has gone through the complete training loop of large-scale AIGC and multimodal interactive models; as early as 2023, he led a To C AI application called "MiaoHua" that gained millions of users in its first month.

An informed source once told us that at the time, within SenseTime's visual domain, there were no more than five people with both algorithm and project experience — Liu, who led a hundred-person foundation model team with over 4,000 GPUs in hand, was one of them.

And much of the stir Vivix has created in investment circles stems from the direction Liu chose for his startup.

In 2024, as the large language model narrative cooled, OpenAI's Sora was stirring up a wave of video generation models on the other side. The route Liu bet on carried the market's hopes for the next form of multimodal models.

At the time, most video model vendors were all-in on improving video quality, longer duration, and stronger physical consistency; Liu, who was then leading SenseTime's entire research team, had much earlier set generation speed and model interactivity as independent technical goals, exploring the use of consistency models to reduce video generation sampling steps.

Liu believes that interactive content is an entirely new content category born in the AI era, and if a model can only do audio and video generation, it's merely reducing costs and improving efficiency for production pipelines built around content formats from the previous era. The origin of his entrepreneurship is this — in the foreseeable future, models have the opportunity to reach an inflection point where interactive content becomes "consumable."

All of Vivix's model training revolves around this goal. "If we're going to build a foundational model for interactive content, first, generation needs to be fast enough, and second, it needs to be cheap enough," Liu says.

This also explains why at his first external communication meeting, the benchmark Liu showed the world wasn't some metric of a top-tier model, but rather YouTube's CDN bandwidth cost — the baseline cost threshold for this generation of content platforms.

In July, Vivix will launch two real-time interactive models, W1 and A1 — W1 is a foundation model integrating understanding, interaction, and generation; A1 is a real-time interaction model centered on character interaction, targeting specific application scenarios like digital humans and character interaction.

The real-time generation cost of the W1 model is already approaching YouTube's CDN bandwidth cost of approximately $0.20 per hour, and compared to current leading video generation models, the generation cost for real-time interactive content can be one to two orders of magnitude lower.

In terms of speed, when a user performs an action in an interactive scenario, the model can generate the next frame within 300-600 milliseconds (TTFF, Time to First Frame). This enables the model to truly achieve the next step beyond personalized distribution — personalized generation.

△ Vivix-W1 demonstrating real-time directing capabilities,

supporting real-time story direction changes through text, voice, and touch input

Vivix's initial chosen scenarios are gaming and livestreaming. This stems from the characteristics of these two scenarios: gaming is many-to-many interaction, livestreaming is one-to-many distribution, with the highest per-token consumption density and higher ceilings for GMV and customer unit price.

Real-time interaction further opens up possibilities for these scenarios. In demos co-created with game studios, players can already use prompts in real-time to control plot progression, truly turning a story into an open-world setting — Liu considers a sense of story as one of the most core metrics for evaluating models. Content that can be interacted with and that creates a sense of participation is the biggest core of next-generation content.

△ Vivix-A1 demonstrating action performance, interaction with objects and environment

And to match this new content format, all existing "infrastructure" needs to be rebuilt. One fundamental reason is that "not everything in the world can be expressed through language, images, and video," Liu says. "We need to throw the raw data of human interaction at the model and let it understand on its own."

In the 19 months since founding, from front-end data collection and data pipeline construction to infrastructure and models, the Vivix team has rebuilt nearly everything. Current team size has reached 150-160 people.

Compared to the ambitious journey Liu has embarked on, he is not an entrepreneur who loves to talk about vision and values, and he's also afraid of people crudely understanding real-time interaction as yet another world model. "I just think world models are too sacred; we're not worthy of discussing them yet," he says.

In Liu's plan, the W1 launching in July is a medium-scale version, with the core goal of validating that the integration of understanding, interaction, and generation can work and can generalize. In Q4 this year, Vivix will also complete training of W2 — with two orders of magnitude more data than W1, covering all imaginable natural video interaction scenarios.

It can be said that Vivix's ultimate goal remains relatively long-term and extremely difficult. On one hand, real-time interaction places higher demands on compute and infrastructure — humans consume language tokens at 5-10 characters per second, while visual content token volume can be thousands to tens of thousands. On the other hand, while real-time interaction has become an area of exploration consensus for next-generation model vendors, the major technical curve has not yet converged and requires long-term investment.

But Liu feels this has instead entered his comfort zone. "Consensus means mediocrity," he says. From childhood to now, whether it was giving up PhD admission to join SenseTime's startup, or continuously internal-startup-ing within SenseTime to the present, the vast majority of his achievements have come from non-consensus.

The following is Vivix founder Yu Liu's account, organized by Intelligent Emergence:

Since our founding, we've built several product prototypes — 7Verse, TipTap, and now we're on our fifth product prototype. All of these products were designed to test one thing: what content formats and interaction methods are truly consumable.

This is a first-principles question in this field. Every company working on AI entertainment, AI content, and interactive content must answer: what kind of AI content would people rather consume than Douyin's content?

We underestimated this question at first, thinking that if we defined interactive content, it would just work — build it and people would like it. After actually building several prototypes, we discovered this is incredibly difficult.

Defining interaction is hard, figuring out what makes content consumable is hard, and both of these difficulties ultimately come back to model capability. So products aren't our top priority. For the foreseeable future, we're still model-first at our core, with the main business being selling model APIs.

Building these products was also an attempt to give the market some direction. When we built Sofia, the first digital human capable of real-world interaction, we found that many otome game companies and AI companion companies wanted to do this too.

They use our API, and the front-end harness takes two or three days to complete — using Claude Code for the entire front-end and data system. What customers accumulate is their own harness — how to sell this product, how to set personality, how to tune game mechanics and numerical systems. The underlying engine is ours.

Before starting up, I was responsible for training all of SenseTime's underlying foundation models. SenseTime's most famous business line was facial recognition — we had about thirty to forty facial recognition business lines: subway face-swiping, ID verification, phone face unlock, passive security. We found that even though all were facial recognition tasks, every细分场景 needed a separately trained model. If you took all the data and trained one model, its accuracy on every scenario would drop.

But this generation of large models is completely different. The self-attention mechanism in the Transformer architecture gives models natural cross-domain generalization capability — the more data, the more diverse the scenarios, the stronger it gets. One GPT model can do translation, companionship, Q&A, coding, and so on, with near-zero marginal cost between these scenarios.

In the first half of 2025, we found more than ten data sources analogous to next-token prediction for LLMs — we call it next-reaction prediction. It's actually a triplet combination, including data across several dimensions: what happened in the past (understanding); how humans interact (interaction); what happens next second (generation).

Games are naturally interactive content — in all game screen recordings and simulation environments, replays can reconstruct the entire game from start to finish just by recording initial states and everyone's operations. This data naturally shows how people interact with content, why they interact, and what changes in content after interaction.

We've already carefully filtered tens of millions of hours of multimodal data, part of which has interactive attributes. Leading general video generation corpora are roughly at this scale too.

This also affects my thinking about how scaling laws work for multimodal interaction data.

Back in the day when we did machine learning training, people needed to label ground truth — for facial recognition, label who this person is; for autonomous driving, draw bounding boxes to mark where cars are.

In the LLM era, the reason scaling laws exist is because next-token prediction in natural language is naturally self-supervised "labeling" that requires no human annotation. In the interactive intelligence domain, we've also found this naturally self-supervised triplet structure of "understanding-interaction-generation," so there exists a scalable data paradigm.

If you look at demand from C-end users, first, they have difficulty paying for tokens; second, their patience for content generation speed is limited. So if we're going to build a foundational model for interactive content, it must have two characteristics — fast enough, cheap enough.

Cost, speed, and quality: these three form an impossible triangle. What we're comparing isn't who has lower cost, who's faster, or who has better image quality, but rather in a four-dimensional space — cost, speed, quality, and content consumability combined — seeing whose model and system pushes furthest toward the upper right.

For true To C content products to emerge in the AI era, cost must break through the cost barriers of current tool-type models. Only when costs can reach the level of what it costs us to browse Douyin or watch YouTube can real-time interactivity become possible.

Our model inference costs are already very close to this state.

What we benchmark against is YouTube. When users watch videos on YouTube, the cost they pay for global network CDN and internet bandwidth is roughly $0.20 per hour.

From March 2025 to April this year, Vivix's model inference costs dropped by two orders of magnitude, and this includes not just generation but also understanding of video and multimodal interaction.

On speed: from day one, the core metric I set for the infrastructure team was TTFR (Time to First Reaction) — how long after a person gives a reaction (say, writing a prompt or tapping on screen) does it generate corresponding reactive content.

For model inference TTFF (Time to First Frame), we need roughly 300 to 600 milliseconds, and including network communication and full pipeline, TTFR is about 1 second. By contrast, leading video models based on a prompt need inference times of hundreds to thousands of seconds to generate video content.

Considering this is a model with approximately 30B effective activated parameters and VAE compression ratio of only 8-8-4, achieving this latency while integrating understanding, interaction, and generation is extremely difficult.

In July, we'll release W1. W1 is a medium-scale model, with the core goal of validating that the integration of understanding, interaction, and generation can work and can generalize — somewhat like GPT-1 and GPT-2. Last month we already started scaling W2; W2 has two orders of magnitude more data than W1. We plan to complete W2 training and release it externally in Q4.

From W1 to W2, what we're scaling is data; from W2 to W3, what we're scaling is model architecture and MoE (Mixture of Experts). For us, right now we're in the data scaling phase — feed the model to saturation first, then scale parameters and framework. Every category of large model — language models, video generation models, audio models, VLM, VLA, including our understanding-interaction-generation integrated model — needs to go through this phase.

What's currently more missing is: how to incorporate all modalities through which humans naturally interact with models into the model's native learning.

The intelligence ceiling of models shouldn't be constrained by human natural language. Understanding of how humans interact with content also shouldn't first be converted to text. My current approach is to throw the raw data of human-model interaction at the model and let it understand on its own.

I wanted to do this inside SenseTime back in 2022, and it hasn't changed until now.

Many investors ask me every first meeting: how do you compete with Seedance? I say, why would I compete with Seedance?

Many domestic companies doing video tools and video models, having chosen this track, inevitably have to compete with ByteDance. But first, it's hard to surpass their model capabilities; second, you can't beat big companies on customer acquisition models and ecosystems — it's very hard to win.

All companies doing video generation models, like those for short dramas or short video creation, are reducing costs and improving efficiency for production of content formats from the mobile internet era. Anything starting from cost reduction and efficiency improvement, in industries where China has advantages, will ultimately have very low profit margins.

Cross-track capability is dimensional reduction strike, not direct competition. It's like using expert systems for RAG versus using large language models for knowledge base retrieval — both do knowledge bases, but the technical paradigm has changed, there's no competition, people will just migrate over.

Essentially, we're creating new incremental value. Right now everyone does short dramas offline; very few do real-time interactive short dramas because there are no corresponding models, and costs are high. In the AI era, only the emergence of new models can spawn new application scenarios. So everything we do is continuously adding capabilities at the model layer, then pushing down to what the corresponding infrastructure should look like and what application scenarios can do.

Big companies will certainly also do interactive content, but the way, motivation, and patience are completely different from startups. Large company decision-making mechanisms lead them to pursue probability endlessly — if something has sufficiently large probability, they invest sufficiently large resources.

The core of startups is pursuit of odds. My own understanding of entrepreneurship is: people who pursue odds find directions with high odds and probability advantages brought by cognitive gaps.

And before this probability advantage grows to where big companies are willing to invest sufficient resources, you need to form your own data flywheel — that is, sufficiently large driving force and sufficiently large moat. Otherwise, startups with no unique cognition, competing with big companies on win rate at the same probability, have no chance.

In the previous era, Adobe could dominate because product functionality was sufficiently complex — I'm a photography enthusiast, but I've never clicked 70% of Photoshop's buttons; yet precisely because of this, all workflows were completed there, with no need to go to other products.

But in the AI era, the barrier is no longer product complexity but model generalization capability. All interaction becomes natural language; users don't need to learn what Gaussian blur or masks mean — one sentence and you get results.

Why was commercialization so poor in the AI 1.0 era? The reason was poor technical generalization — you had to do high-unit-price, high-margin things to subsidize R&D. I'd build a model for this company, and only they could use it; I couldn't give it to another because performance would drop.

The underlying variable of this generation's real-time interaction technology is that model generalization is strong enough, meaning business model changes — the marginal R&D cost of replication capability is zero. No longer do you need to rely on one customer's high margin to subsidize customized model R&D costs for another customer. What we should earn is the added value brought by model-based scale.

And looking further into the future, companies with model training capabilities will definitely be more than one, model capabilities will eventually converge, and homogeneous model business models will still trend toward utility-style competition, with the market entering zero-sum competition and transparent channels, and gross margins declining. So we need to create added value in scenarios and channels during this process.

AI technology innovation and breakthroughs cannot be planned; you need a loose environment to stimulate creativity.

But in this era, to reach the first tier of models, you can't rely on just one genius idea — you must have very strong systematic operational capability: very strong data infrastructure, training infrastructure, and data-algorithm capabilities. The devil of ultra-large-scale training tasks is in the details. Give the same 10,000 GPUs, 1B raw data to different teams, and results could be vastly different.

For example, why Seedance is so impressive — I personally feel there are two points. First, its legion operational capability is strong enough — data, models, algorithms may not have too many fancy innovations, but every detail is done well, and the model team works hard enough; second, automatic storyboarding and prompt-following capability are done well, greatly improving product usability.

I met quite a few ByteDance people in the past two years, and was always curious how ByteDance could achieve this — no matter who changes, the entire team continues to operate at maximum combat effectiveness.

Why does this large-legion form of operation work today? Because at that time in the video generation track, DiT and MoE were already consensus — you didn't need to do overly fancy innovations; instead you needed legion operations to get every detail right step by step. This is where organizational advantage can play out.

Today, using AI to produce content is still in early stages. The entire video generation track's ARR has increased tenfold or more compared to last year, and in the coming years this track can be expected to grow by orders of magnitude. Our bet today is using models to directly change content formats — this track is earlier, more unknown, and more interesting.

Many people ask me how I view world models — it seems like everyone's starting to do it, so it's become consensus? But I don't think so. Consensus is when your cognition aligns with everyone else's, but people's motivations for doing this are clearly still inconsistent.

On a track that hasn't yet been established, it's wrong to compete on a wrong consensus metric. Like with digital humans, everyone's competing on who has better lip-syncing, whose hands look more natural, but I don't think this is very fundamental. I like watching Yonghao Luo's livestreams because he's witty; I like listening to Zhang Xuefeng talk about college entrance exams because he can explain the gray adult world to me. These are what consumption is about.

In my own growth, the positive feedback I've received has never come from consensus. Basically all my work has required entrepreneurial spirit — starting a business while studying at CUHK, doing research at Microsoft, doing research at Google Brain, going from intern to research director at SenseTime, then to leading a business unit — all were zero-to-one things.

In 2015, I applied to a pretty good school for PhD, visa was already processed, but Professor Tang Xiaoou hoped I'd stay to start a business together. I was 20 that year, chose to give up the PhD, and stayed at SenseTime. Looking back today, I don't regret it at all.

Consensus means safety, but also means mediocrity. I've been someone who doesn't pursue safety since I was young — safety gives me extreme insecurity.

Someone asked if I'd worry about people defining us as a world model company — I actually mind quite a bit.

I don't really want to talk about world models. On one hand, the term has been overused with very泛滥 definitions. But the deeper reason is, I think it's too sacred, and we're not worthy of discussing it yet.

One view I hold is that language-centric multimodal models will not become things that can truly represent this world.

First, training corpora for language models all come from the internet — things written by those who once had the ability to leave their voices and records on the internet. This corpora is inherently biased.

I'm personally interested in neuroscience and cutting-edge high-energy physics. If ChatGPT could truly help me understand the cutting-edge cognition in these fields, it would be a world model. But I find it extremely biased.

Furthermore, assuming corpora had no bias, every line a person writes is essentially electrical signals produced by neural circuits in the brain, which become text when they reach the language center. And these electrical signals stem from everything they've seen, heard, and thought in their life — essentially their sensory system. Vision is essentially electromagnetic waves; hearing, smell, and touch ultimately also come down to electromagnetic force. This is merely one of the four known fundamental forces that constitute this world, and cannot represent the entire world.

All human cognition is some abstraction of perception. Human definitions of "world models" also cannot transcend humans' own perception of "the world." For example, on the electromagnetic spectrum, the human eye can only perceive visible light, so we have a media format called "video" that records the visible light band; and being able to predict next-frame "video" content based on action is defined as a "video world model" — this is narrow. Human-condensed text originates from such narrow observation of the world.**

If humans hadn't discovered X-rays, and relied only on videos recorded with visible light to say "this can represent the world," wouldn't you find that absurd? Perhaps in the future there will be Y-rays, Z-rays, or a new fifth fundamental force that can bring us more essential physical understanding of the world, then wouldn't "physical world models" based on visible light and such seem somewhat narrow?

"World" is itself defined by humans based on limited observation. And humans are narrow; things defined by humans must be even narrower than humans.

Having thought this through, I actually feel relieved. If we understand world models from a narrow angle, it's essentially functionalism: it can help humans predict what happens in the next frame, can let humans see content that looks somewhat like what people filmed. This "world" is a world to satisfy human practicality, not the real objective world.**

After AI agents save humans massive amounts of time and produce massive productivity, the time humans have left should be filled with spiritual consumption in more efficient ways. If I can achieve this step, I'll be satisfied.


Moonshot AI K3 Release: A Sleepless Night, We Witness the Fulfillment of Long-Termism


BlueRun Ventures' Jui Chan in Conversation with BAAI's Wang Zhongyuan, Galaxy Universal's He Wang, and ModelBest's Li Dahai: Long-Term Value in the Large Model Era and the Next Curve


BlueRun Ventures Leads Round, SiClink Focuses on Visual Reconstruction, Exploring New Directions in Brain-Computer Interfaces | BlueRun Ventures Family