Vivix: Starting from a Place Where No One Believed
Founded in early 2025, Vivix is a portfolio company of Monolith. For the past 19 months, the company has remained almost entirely silent. **Beneath the surface, it has completed five funding rounds at a valuation of $1.32 billion.**



Vivix was founded in early 2025 as a portfolio company of Monolith. For the past 19 months, the company has barely made a sound in public. Beneath the surface, it has completed five funding rounds at a $1.32 billion valuation.
This month, Vivix made its official debut: releasing two real-time interactive models — W1, a foundation model for interactive storytelling, and A1, for interactive AI characters — along with API access. This marks the first time founder Yu Liu has systematically explained to the outside world what the company is actually building.
Born in 1995, Liu did visual AI research at The Chinese University of Hong Kong's MMLab in his early years. At age 20, he abandoned a PhD offer with visa already in hand to join SenseTime in its second year. At 26, he became one of the youngest division general managers in SenseTime's history. At the end of 2024, he left SenseTime to start his own company, choosing a path that almost no one believed in at the time: real-time interactive multimodal models.
For the past three-plus years, the entire industry has been searching for the next consumer application. Our view is: only innovation at the interaction layer will lead to new application scenarios and new products. And in the current technology cycle, innovation in interaction is fundamentally a model problem — and the essence of a model problem is the balance between real-time capability, comprehension, and generation. As for what the final application will actually look like, no one knows.
The most significant change at Vivix over the past year is that its model and infrastructure capabilities have leveled up every two to three months: first putting comprehension, interaction, and generation into a single model, and now beginning to scale data and frameworks. Liu is also an unusual founder — his conviction, aesthetic sensibility, and way of thinking through problems have all left a deep impression on us. This is a hard path, but we're excited to walk it with Vivix.
Based on Liu's recent external communications, we've compiled this first-person account. We hope it helps you understand the company.
When No One Believed
Why did I start a company? Because I had discussed this internally at SenseTime, and in the end there was no clear project approval. So I thought, I'll just go do it myself, spend some of my own money to build prototypes.
It was the end of 2024. Model entrepreneurship had already reached a relatively cold state. The consensus was: the model war dominated by LLMs, the zero-to-one phase, was largely over — all the noise had shifted to AI applications.
Technically, at that time the best domestic video models took 5 to 10 minutes to generate 5 seconds of content. The viral case everyone talked about was "bringing old photos to life": making a black-and-white photo move slightly, with no sound, and that was already considered amazing.
In an era like that, just "generation" alone was already so difficult and so expensive. And what we wanted to do, on top of generation, required comprehension — understanding what happened in past content; and interaction — understanding what your voice, touch, gestures were trying to do in this moment; and then responding in the very next frame. Three intelligence tasks, one model, real-time, and cheap.
So it was completely normal that people didn't believe. When the first-round investors bet on us, I was actually quite surprised: there really were people willing to bet on this.
But that period was tough for me. Many investors I met asked: when will your product launch? But for what we're doing, the pragmatic sequence is: first build the model, then with model capabilities in place, build infrastructure to get costs and experience right, then build the product. Just like first there was GPT, then the chatbot, then later companions and coding.
Later when I talked to Xi Cao, he said right away: You don't need to rush the product. I remember it very clearly. We talked late into the night in our very small office at the time. Our first model was already out, with tens of thousands of test users in North America. I showed him the backend data: how users liked to play with interactive content, which directions the model should go.
At this level, we had resonance — our views on things, our starting points and ways of thinking, were intuitively aligned. About ten minutes after talking with him, the financing closed quickly. I felt like I had met a kindred spirit of an institution.
Killed Prototypes, Gained Convictions
Our entire team comes from technical, research, and infrastructure backgrounds. When I started the company, I was clear about what we needed to validate: at the level of "what kind of model do users actually need," we needed more understanding. For a mature researcher, defining next-generation model architecture is easy; defining what model has market demand and can form a business model is hard. This can only come from high-frequency market feedback.
So my thinking was very basic: whenever we built a model, wrap it in a shell and give it to users — the core was collecting feedback. Every killed prototype was substantial positive feedback for me. Because every problem we encountered was new, and would become part of our conviction.
First conviction: interaction is not consumption. The first prototype tested interaction itself — touch, gestures, camera, gyroscope. As long as there were different sensors, there were different interaction methods. But which ones did users actually like, which were they willing to pay for? No way to guess, only to test. After testing, we quickly killed it, because we found that in that era, the consumption quality of AI-generated content was far below UGC. People would rather use Douyin than play with this.
Second conviction: what users consume is story. We spent two to three months researching one question — what is consumption power? We surveyed all the content formats that the internet era ultimately converged on: short videos, short dramas, YouTube and Netflix series.
The conclusion: what users are truly willing to consume is story, not attention-grabbing. Short videos may only be 15 seconds, but they tell a complete story: the first three seconds have a hook that grabs you, the next ten-plus seconds have constant twists. These are the foundational elements of good storytelling, while AI doing some special effects or transitions falls far short on story.
Later, something confirmed these judgments. Last National Day, OpenAI's Sora2 App came out, exactly retracing the path of our first prototype. At the time everyone asked me: didn't you guys do something similar, why did you stop?
My judgment at the time was that this thing wouldn't work, but after all it was OpenAI, so I couldn't go around preaching. It ultimately proved that the problem was exactly this: it was still making a tool, just a more mass-market tool. What users want to consume isn't the act of generating images itself, but interesting content.
Third conviction: content must be human-centric. Among the converged content formats, seventy to eighty percent of frames revolve around people: dialogue, talking-head videos, livestreams. So we branched out a human-centric model line, which is this A1 release.
The digital human concept has been around for nearly a decade, but we passed on that direction from day one: digital humans that can only move their lips and wave their hands have no story, and don't add much to content consumption. Our definition of A1 is that the person must be able to have stories happen to them in a scene.
In a live commerce scenario, he really holds the product, points at it, and talks to you about it. In a gaming scenario, he really can walk in the environment and interact with it.
Fourth conviction: interactive content is still in a very early stage today. Breakthroughs in experience depend on good content creators, and also on fundamental breakthroughs in model capabilities. Otherwise interactive content is just拼接 of offline content, not truly scalable experiences. For the foreseeable future we are committed to building the best interactive models, hoping to serve more creators interested in interactive content.
A Dozen Data Sources No One Noticed
Why did language models succeed? Because all text written by humans on the internet is naturally their training material, requiring no human annotation.
But what we need our model to learn is something else: how people interact with content, and how content changes after interaction. This "textbook" doesn't exist ready-made in the world. In the first half of 2025, we found a dozen data sources that no one had noticed — actually a triplet combination, including several dimensions of data: what happened in the past (comprehension); how humans interacted (interaction); what happens in the next second (generation).
Games are naturally interactive content. In all game recordings and simulation environments, replay only requires recording the initial state and everyone's operations to reconstruct the entire game from beginning to end. This data naturally demonstrates how people interact with content, why they interact, and how content changes after interaction.
How humans interact with content shouldn't first be translated into text. Our approach is to feed this raw, native data directly to the model and let it understand on its own. The intelligence ceiling of a model shouldn't be bounded by human natural language.
Our current data volume is roughly comparable to the corpus scale of the best video generation models, with some portion being content with interaction.
This W1 release is a medium-scale version, core to validating that comprehension-interaction-generation integration can run and generalize; W2, with data volume scaled up by two orders of magnitude, will release in Q4. Feed the model first, then scale parameters and frameworks.
Costs Need to Hit Douyin Levels
Why did no one do real-time video interaction before?
Do the math and it's clear. Humans consume text at a dozen or so tokens per second; visual content can be thousands to tens of thousands of tokens. Doing real-time at this density, most teams thought it wasn't realistic.
And consumer users have two typical characteristics: unwilling to pay for content, and no patience.
This means the underlying model for interactive content must be both fast and cheap. Fast is actually easy — just throw GPUs at it. Serve one user with a thousand cards, any model can be "near real-time," but that's meaningless.
From day one, the core metric I set for our infrastructure team wasn't any model metric — it was cost benchmarked against YouTube: for every hour we watch YouTube, Google pays roughly $0.20 for global CDN and bandwidth. This is the cost floor of internet-era content platforms. Only when costs drop to the level of us scrolling Douyin or watching YouTube can real-time, interactive content be mass-consumed.
Our curve: from March 2025 to this April, costs dropped by two orders of magnitude in one year. And this number doesn't just include generation — it includes comprehension of past video and multi-turn interaction. It's the cost of the entire system.
On speed, we have an internal metric called TTFF (Time to First Frame): after a person gives a reaction, writes a prompt, or taps the screen, how long until they see the next frame of content.
We're now at 300 to 600 milliseconds — and this is on a unified comprehension-interaction-generation model with 30B activated parameters.
Talking about latency without talking about parameter count is耍流氓.
Startup Opportunity Lies in Conviction Gaps
Many people ask me how I view competition with the big tech companies? The big companies will definitely do interactive content.
But the decision-making mechanism of big companies means their pursuit of probability is endless: however probable something is, that's how many resources they invest. This is the most efficient way for organizations with sufficient resources.
Startups don't have that many resources. The old saying goes that the barefoot aren't afraid of those in shoes, but being barefoot isn't about gambling. What startups truly need to pursue is conviction gaps: you must find a point where your conviction differs from that of big company managers, a point where your probability is significantly higher than theirs.
What is 10% for them is 30% for you. 30% still isn't high, but you're willing to bet all your resources, time, and energy on it. Meanwhile, big company managers may have a hundred 10% probability things on their plate, each getting only one percent of their time, a one-hour weekly meeting asking the innovation team "what progress have you made." Their conviction iteration is far from sufficient.
Only at such points do startups have a chance: before the probability rises to where big companies are willing to enter and compete with you, build sufficient momentum and sufficient moat. Otherwise, startups without unique conviction, competing with big companies on the same probability, have no chance at all.
As for the present, so far we have no real competitors. The reason is simple: this track hasn't been proven to work yet — how can there be competition?
At this point in time, converging on false consensus isn't meaningful. Defining success for what we've defined — from PMF, from technology, from models — is the only thing we need to consider.
Also, many people ask if I worry about being defined as a world model company. I mind it quite a bit. I don't really want to talk about world models. On one hand, the term has been overused to the point of dilution; more fundamentally, I think it's too sacred, and we're not worthy of discussing it yet.
My essential definition of a world model is: it must be able to represent, even reconstruct, the states and dynamics of motion of the physical world. And all human understanding of the world is based on our own senses: vision is electromagnetic waves; touch is electrical signals from neurons, the essence of pressure sensing is repulsive force between atoms; smell is chemistry, and the essence of chemistry is still electromagnetic force.
There may be trillions of modalities in this world; humans can only perceive a few of them. All the corpora, videos, and audio we collect were recorded because humans can see and hear.
Training a model on these things and saying it can represent the dynamics of the world — I find that a bit funny. What we want to do is actually more pragmatic: use AI's capabilities to create value for human spiritual consumption.
The Number One Has No Exit
For my first six years at SenseTime, I managed the most upstream model team in the research institute; for the next three to four years I led a division, responsible for revenue, gross margin, and sales results.
The management methodologies for research teams and delivery teams are vastly different, some points are actually opposite: research teams need enormous tolerance for failure, with ample space and delegation; this model doesn't work for delivery teams — humans are lazy, if you're infinitely tolerant, standards will gradually drop.
I think my DNA is in the former. So when I started the company, I was indeed managing the entire company with an innovation team model. The benefit was that we did well on technology and prototype iteration; the problem was that on delivery lines where results mattered, my management was insufficient.
Entrepreneurship is the fastest training for a number one, bar none. As long as you're not number one, there's always someone above to cover for you. If you make a mistake, at worst your bonus drops a bit.
The number one has no exit.
When you've already given up things that ordinary people can't imagine for this company, you won't give excessive tolerance to things that shouldn't be tolerated.
Judgments on resources, time, and people become much more decisive.
I feel like in management, every month I'm different from the last month, constantly evolving. This month it's judgment on people, next month it's organizational model. The underlying driver is the same: you're number one now, you've bet everything.
Some people ask, isn't starting a company a very hard decision, with such high opportunity cost. But for me, entrepreneurship isn't a choice, it's a state of being — it's the easiest decision.
Currently, we have a major breakthrough in models and infrastructure every two to three months — not incremental from 50% to 70%, but leaps from "no one believed real-time generation was possible" to "real-time generation works."
We've consistently delivered things no one has seen before, and we hope to keep doing so.


