How many steps does it take to become a "down-to-earth" AI founder? From Google X researcher to building a product that hit $1M ARR in six months | A conversation with Vozo founder Changyin Zhou
Changyin Zhou's Journey Fusing Technology and Business
Not long ago, we used 20 questions to map out the past year of AI video models with Barkley, the sole product manager at Luma.ai[1].
For this podcast episode, we invited Changyin Zhou, founder of Vozo — a one-click video remixing tool — to share the story of building Vozo: it can redub videos, translate them, and re-edit them. At launch, it topped Product Hunt for three consecutive days, and hit $1 million ARR in six months.

In our conversation, we also used Vozo as a thread to explore Zhou's previous work at the famed Google X lab, his first startup in Silicon Valley, and how he began his entrepreneurial journey back in China with the decidedly down-to-earth "live streaming machine."
We had assumed that walking away from past ways of working would take enormous courage. But Changyin told us that, compared to sunk costs, he cares more about whether he's doing something no one has done before.
He also shared observations about the work habits of exceptionally smart people he's known — we hope you'll find them valuable.
Listen on WeChat:
Listen on Xiaoyuzhou:


Vozo AI: An Innovative Journey in Video Translation and Remixing
🚥 Ronghui
Last episode we reviewed the progress in video models since Sora's release. Today we're talking with Changyin Zhou, founder of Vozo[2], about his very specific entrepreneurial story and personal journey.
Vozo has been described as "a one-click AI tool for remixing short videos." It can redub videos, translate them, and re-edit them. It topped Product Hunt for three straight days at launch and reached $1 million ARR in six months. From what we understand, Vozo's development had several interesting inflection points, along with Changyin's corresponding thinking at each stage.
Today we've invited Changyin to walk us through these stories in detail. Changyin, please say hello, introduce yourself, and tell us about Vozo.
👦🏻 Changyin Zhou
Thanks Ronghui. I'm Changyin Zhou — you can call me Changyin. I'm the founder and CEO of Vozo AI.
🚥 Ronghui
Could you walk us through what Vozo specifically does, and how these features were developed step by step?
There are quite a few AI video tools out there — why did you choose this particular direction?
👦🏻 Changyin Zhou
Vozo went through a long "gestation" period internally. In 2021, our team returned to China from the US and decided to focus on the concept of "video expression freedom." We believed this was something very much worth doing.
Starting in 2021, we launched several products — some successful, some less so. From 2022, we began exploring generative AI R&D. This was a two-pronged approach — starting from user needs on one hand, and from technology development on the other. By 2023, we had some ideas and found the intersection where R&D and demand collided. We began internally incubating and screening several concepts. In 2024, when we felt satisfied with the product ourselves, we officially launched Vozo.
Vozo's positioning evolved several times, but the core has always been helping ordinary people, not professional video editors. "Ordinary people" here covers a wide range — teachers, project managers, marketing managers, and so on. They occasionally need to make videos, but have limited professional skills or need to outsource. Our vision is to enable all ordinary people to express themselves through video. This is actually a very big ambition. Our initial generative AI model was quite radical, similar to many of the generative visual large models people see today.
But at the end of 2023, we pivoted to focus on real-world scenarios and genuinely solving user problems. So when we launched in July 2024 and first campaigned on Product Hunt, we defined the feature as "Vozo Rewrite." We appropriately lowered the difficulty — instead of generating videos from scratch, we changed the story based on existing video.
This function applies to multiple scenarios. One is using existing high-quality video footage, such as movie clips, to tell your brand story or viewpoint. Another is easily converting existing videos — say, a Thanksgiving promotion — into a Christmas promotion.
I think Vozo Rewrite brought quite significant changes to video editing. Traditional editing requires cutting and audio-visual processing. With Vozo Rewrite, you simply use a prompt like "Please convert this video to Spanish" or "Make the video more exciting / more fun," and the transformation happens. This was the core function of our initial launch.
After more than half a year of evolution, Vozo's feature set has become much broader. It's worth noting that after Vozo's first release on July 20, we iterated multiple times. By November, the product underwent a major transformation with the launch of our new feature "Vozo Translate."
Translate is actually an extension of Rewrite, because we discovered that a large number of users were using Rewrite for translation purposes. So we iterated internally for a long time, inviting users who needed translation to try it out and provide feedback, gradually refining the translation feature.
Currently, most of our users have very high renewal rates. Counting from July 2024, although we started relatively late, by January-February 2025 the product form was quite mature. That's the development journey of Vozo — our company's latest product.

From Idea to Market: Vozo's Feature Iteration and Commercialization
🚥 Koji
I first encountered Vozo in the Crossing membership group. At the time, Vanessa (a PM at ByteDance who also appeared on the Crossing podcast episode "AI Product Manager Guide") recommended the product to everyone.
Her recommendations are consistently high-quality, so I paid extra attention. What left a deep impression was seeing multiple short videos generated with Vozo going viral simultaneously — these videos cleverly adapted classic movie scenes into comedy skits. The visuals remained unchanged, but the protagonist's lip movements and tone were completely different.
I remember one featuring Leonardo DiCaprio's classic scene from The Wolf of Wall Street. He maintained the same impassioned delivery from the original, but the content became mundane daily trivialities — the stark contrast was striking. Beyond that, scenes from Titanic, Harry Potter, and various other classic films were creatively "remixed," feeling refreshingly novel. This should have been Vozo's first wave of breaking into mainstream awareness.
By November, Vozo hit the trending charts again, even becoming Product Hunt's monthly #1. This time it was because of the newly launched translation feature, which could perfectly convert video content from one language to another — the quality was so outstanding it won widespread acclaim.
Apart from these two major product highlights, is there anything else you'd add, Changyin?
👦🏻 Changyin Zhou
Currently "translation" is indeed our most frequently used feature.
As mentioned earlier, the Translate feature officially launched in October. In the three months after November, we successively developed two additional product features.
These three features all emerged gradually. We initially developed Rewrite, then discovered most users were using it for translation purposes, so we decided to deepen the Translate feature — a process that took considerable time.
After refining Translation, we noticed some users didn't need full translation services at all — they only wanted to use our lip-sync technology. Based on this demand, we deepened development of the Lip Sync feature, which has now become one of our important functions.
Interestingly, after launching Lip Sync, some users raised a new request — they wanted to do lip sync on photos rather than videos. Initially we hesitated at this demand, thinking there were already similar Photo Lip Sync tools on the market. But after testing and analyzing various existing tools, we understood what dissatisfied users about current solutions. So we redeveloped the Photo Lip Sync feature and launched it around January. After this feature went live, user growth was rapid, which also proved our results were genuinely satisfying.
A brief teaser: we'll have something bigger releasing in March, but it's confidential for now (laughs).
🚥 Koji
Will this bigger thing bring visual shock like the original Rewrite did, or will it be like Translate or Lip Sync — functionally better than competitors?
👦🏻 Changyin Zhou
Actually both. More of a feature based on existing demand, but we're doing it quite differently.
🚥 Koji
Do you think Vozo would have struggled to gain market attention if it had initially pushed translation or lip sync as its main features?
And that choosing to launch an unprecedented innovative feature precisely sparked user curiosity, enabling short videos to break through and spread, successfully bringing Vozo into a broader user base?
Now that you've tasted the benefits of this innovative marketing approach, will you continue pursuing product and marketing-level innovation? Like the original Rewrite feature, or how Pika now releases new effects monthly. Does this represent your chosen path going forward?
👦🏻 Changyin Zhou
Yes, I think the path really matters:
What your first feature is and what your first impression is — that's extremely important.
After all, I believe in this GenAI era, innovation is essentially the primary marketing tool. So you definitely don't want people to see you as a "me too." And "me too" is also hard to justify internally — for an innovation-driven team, it's difficult to keep doing "me too" work because morale will just collapse.
Of course, if your team isn't an innovation team, then it doesn't matter. If you're already a "me too" team, just keep doing that. But in my biased view, "me too" teams simply cannot succeed in today's AI era.
If you're an innovation team, you have to constantly produce new innovations and push them out. But there's a slight paradox I mentioned earlier — we need to capture demand while also having innovation points, so the path becomes crucial.
After launching as an innovative brand, pivoting from Rewrite back into the Translate space allows us to achieve better results than traditional translation. Translation demand is real and widespread. Our strategy is to operate with precision first, then gradually expand our scope.
This path is quite favorable for both technical evolution and market expansion. Not all business models fit this trajectory, but for AI video, we were fortunate to find this path of continuous expansion.
It's worth noting that some companies choose different strategies — quietly developing without revealing any roadmap, then suddenly blowing up. Such cases do exist. We chose to start from user needs and explore AI video application scenarios. That's our current focus.
🚥 Koji
The AI video space has seen many innovative companies emerge, from Pika to Luma, from Heygen to Viggle, OpusClip, and others. Which companies do you particularly admire or follow?
👦🏻 Changyin Zhou
I quite like Heygen — I think they're very focused on what they want to do.
Founder Josh Xu started from the vision of "replacing traditional filming" and has steadfastly pursued this direction since 2021. Despite numerous technical challenges along the way, the team gradually broke through. Regardless of how one evaluates their decision to relocate to the United States, Heygen has truly excelled at product refinement and technical iteration.
My knowledge of other companies is relatively limited. For example, Dzine, which works on image-related products, was founded by another friend of mine, and their product quality is also excellent. Personally, I prefer teams that obsess over product experience — both Dzine and Heygen stand out in this regard.

Breaking Out on Product Hunt: Cold Start Strategy
🚥 Koji
I know Vozo hasn't spent any marketing budget — just two Product Hunt launches — and you've reached $1 million in annual revenue today.
How much do you think Product Hunt launches have actually helped you?
👦🏻 Changyin Zhou
I think they've helped quite a bit, and it needs to be viewed from two angles.
First, I genuinely love Product Hunt. I actually did my first Product Hunt launch back in 2015. The atmosphere was quite different then.
But I think its core value is that when you do a Product Hunt launch, you really have to think about what kind of product you are and how to explain it in one sentence. I think this is actually the most useful thing for product refinement. So I believe this is where Product Hunt's greatest value lies.
What it gave us was a relatively simple cold start. Although the traffic wasn't huge — roughly 1,000 visitors per day — that 1,000 was enough for us to iterate on product-market fit, so it served as a cold start. Completing this through one or two Product Hunt launches — I think that was extremely valuable.
🚥 Koji
Actually, there are tons of tactics for ranking on Product Hunt. I see people canvassing for votes in various groups every day.
For a product to hit #1 on Product Hunt's daily ranking today, how much of it is operational maneuvering versus the product itself needing to be genuinely excellent?
👦🏻 Changyin Zhou
This is exactly what I meant at the beginning — there's a huge difference compared to 2015. None of this existed back then.
Today, I think getting a high ranking on Product Hunt isn't actually that closely tied to the product itself.
If you understand operations and are willing to invest resources in promotion, you can basically push any product to the top. Maybe you can't guarantee #1, but top three shouldn't be a problem. But I think this is just one aspect — getting to #1, #2, or #3 doesn't mean you've truly succeeded on Product Hunt.
Your success depends on whether this Product Hunt activity ultimately produced substantive help for your product's product-market fit (PMF). So I'm not surprised that many teams might rank #1 or #2, yet their product ultimately fails to land.
So, returning to your question — getting a high ranking on Product Hunt is primarily an operations matter. Whether you can generate real commercial value after achieving that high ranking — that's the product-level problem that needs solving.
🚥 Ronghui
Besides Product Hunt, have you done any other kinds of exposure?
👦🏻 Changyin Zhou
We've barely done any additional promotion. There were some opportunities along the way, but we've been relatively restrained. Mainly because Product Hunt already gave us sufficient traffic, so we cherished that six-month window to focus on the product itself.
Another reason is that after the initial traffic came in, we received a lot of user feedback. We felt that further promotion wouldn't make much sense before addressing these feedback issues. The feedback was indeed substantial, but I think we did a few things right: we implemented an Intercom-like customer service system very early on, so users could communicate with us directly on the webpage.
Some of our team members would stay online to respond, understanding what users truly needed, what they didn't need, and what dissatisfied them. Based on this feedback, we continuously iterated the product, releasing roughly one to two versions per week, constantly improving it.
So during that phase, we didn't focus too much on promotion. This approach may not suit every team — some might choose to promote earlier for faster growth. But this is my view:
Whether you promote one month earlier or later actually doesn't matter that much. Ultimately, product-market fit (PMF) is the more important factor.
🚥 Ronghui
Was there a moment when you felt you'd found PMF?
👦🏻 Changyin Zhou
I think it's a feeling, but if we're talking quantitative measures, we focused on two key metrics: user renewal satisfaction and the absolute value — that is, our annual recurring revenue (ARR).
Our approach was fairly straightforward: set a goal of reaching $1 million ARR first. Fortunately, even without additional promotion, we just hit that target. There may have been some luck involved, but it also largely aligned with our initial judgment.
Another key metric is renewal rate, which I believe needs to stay at a reasonable level. For example, if we have 100 paying users and they're satisfied with the product, roughly 80 should choose to stay and continue using it. Of course, this is my personal assessment, since I know about 20% of users might stop due to their own business reasons. So when the renewal rate reaches that level, I consider our product to be passable.
These two aspects combined became our internal goals. The benefit of this approach is that with clear goals, you have more motivation to execute. Focus on one or two goals at each stage rather than doing multiple things simultaneously — like improving the product while also promoting through five different channels. We tried to separate these tasks and push forward more purposefully.

Technical Struggles and Team Transformation: Finding the Right Path
🚥 Koji
You mentioned earlier that Vozo didn't officially launch its first Rewrite-focused version until July 2024 — upload and "remix" existing videos — which quickly went viral upon release.
Compared to many competitors, this timing is indeed somewhat later. I'm curious about this timing choice: was it because the technology only truly matured at this point, or were there other reasons?
👦🏻 Changyin Zhou
I think both factors played a role, but our own internal reasons probably accounted for the larger share. We later did some internal retrospectives.
We actually started very early in the AI video space. When we founded the company in 2021, we were already working in this area, though perhaps more oriented toward traditional computer vision (CV). By 2022, we began researching some deeper-level technologies.
In 2022, we foresaw the potential of generative AI earlier than most companies. So we did something quite unusual for an early-stage company: we established a joint lab with an external, well-known researcher who was also a former professor of mine. We invested significant resources in foundational research and frontier exploration, and at that time we had almost no revenue — a fairly bold decision.
These R&D efforts began showing results around early 2023, particularly as generative video models started emerging. That period was incredibly exciting — we could iterate a new model roughly every two to three weeks. But we made one wrong move at that time — we were doing two things simultaneously: on one hand, promoting our existing product and pursuing revenue; on the other hand, conducting foundational research, believing this might be a crucial opportunity for the future.
At the time we had a theory that seemed sound: pursue both directions simultaneously — building market-facing applications on one side, and cutting-edge research on the other — hoping the two paths would eventually converge.
But looking back now, this was actually a deeply flawed idea for a startup.
By 2023, we found ourselves in an awkward position: the product features we wanted to build couldn't be supported by our foundational models. The research had proceeded along its own trajectory, and while the model results were interesting and exciting, they proved difficult to productize — plagued by all sorts of randomness and strange issues. So for most of 2023, we were basically stuck in a loop — aggressive on research, aggressive on application development, yet the two paths never effectively merged.
🚥 Koji
No synergy — instead, a situation of "you can't help me, I can't help you."
👦🏻 Changyin Zhou
Exactly, incredibly frustrating. The researchers were frustrated too — they'd produce a model and wonder why we couldn't productize it. The product team would say they needed something, and ask why the model couldn't deliver it.
🚥 Koji
Which side won out in the end?
👦🏻 Changyin Zhou
In the end, the application-focused, user-needs-driven direction won out. By October 2023, we issued a PR announcing our model results. While the PR didn't explicitly state this, it effectively signaled that we were no longer pursuing pure research.
We released a multimodal model called "HiveNet," but from that point on, all R&D projects had to originate from product requirements and gain product team approval to proceed.
In theory, we still reserved 20% of resources for the research team to explore directions they found interesting. But after October 2023, all our R&D initiatives started from product needs and were oriented toward solving real problems.
🚥 Koji
But wouldn't researchers leave because of this? Wouldn't they feel this was no longer the research environment that had attracted them?
👦🏻 Changyin Zhou
Actually, no.
After more than a year of development, most researchers wanted to see their work integrated into products. They noticed our other products had substantial user bases, yet their own research struggled to make it into those product lines. This really was a mindset issue among researchers.
We were fortunate to collaborate with several researchers who were particularly focused on application落地. Now, whenever Vozo's user numbers grow and feedback is positive, these researchers feel genuinely happy. I believe this has created a virtuous upward cycle.
🚥 Koji
If I may ask, where is Vozo in terms of funding? What's the team size? And how is it split between research and product?
👦🏻 Changyin Zhou
We're pre-Series A, mainly backed by Linear Capital and HSG. There are also some individual investors, bringing the total to roughly $6–7 million, so our capital efficiency is probably quite high — we've iterated through many products along the way.
🚥 Koji
Only $6 million raised across four years from 2021 to 2025 — that's remarkably efficient.
👦🏻 Changyin Zhou
Starting from 2022 to 2023, some of our earlier products were quite successful and generated revenue, so our financial situation has been relatively healthy. The entire team is currently cash-flow positive, so the pressure isn't that intense. This was something we didn't initially anticipate, but we later realized it has a very positive effect on team morale.
Our team now has over forty people. R&D accounts for roughly 70% or more, so we're very heavily invested in research.
🚥 Koji
I'm curious — how does a team of just over forty people achieve break-even with millions in ARR? Does this mean you have other product lines continuously contributing revenue?
👦🏻 Changyin Zhou
While Vozo is where I invest the most energy, it's not currently our main revenue source.
We previously developed two apps — in China they're called "Shuode Teleprompter" (now renamed "Shuode Camera"), and in overseas markets, "Blink." These apps also aim to help creators produce video more easily, but they're built on more traditional computer vision and natural language processing technology.
These two apps generate roughly $6 million in ARR. It's these products that ensure our current cash flow breaks even. So all of Vozo's ARR is essentially our profit.
🚥 Ronghui
Are you an app factory now?
👦🏻 Changyin Zhou
That's a good question.
Initially we didn't have a clear direction — we simply developed our first app around the concept of "freedom of video expression," based on user needs. Later we realized this app was actually constrained by traditional computer vision methods, so we began developing new products based on generative AI technology.
For a considerable period, we operated both products simultaneously, which was particularly painful for our team — advancing two parallel product lines at once. But over time, we've now found an effective way to integrate them. Before long, you'll see these two products actually merging into a unified product. Their features will be shared with each other, ultimately serving all content creators, marketing managers at various companies, e-commerce practitioners — anyone who needs to tell visual stories through video.
🚥 Ronghui
What method did you find to effectively combine them?
👦🏻 Changyin Zhou
The user overlap between these two products is roughly 20% to 30%.
In terms of positioning, the app side mainly targets C-end users, including KOLs, KOCs, and a small number of SMBs. Vozo, meanwhile, primarily serves enterprise marketing departments and some SMBs, so we have considerable overlap in the SMB segment.
After merging, they'll achieve cross-user referral and feature sharing. We'll establish a unified membership system — users who purchase a Vozo membership can also access features in the app; similarly, app members who add a certain number of credits can use Vozo features. This way, users on both sides will be fully connected, and we're very excited about this.
The merged product will统一 use the name "Vozo," because the entire team prefers this brand name.
🚥 Ronghui
Why the name Vozo?
👦🏻 Changyin Zhou
GPT actually created this name for us — quite interesting, and it left a deep impression on me.
We wanted a short word related to "video" and "voice," because when we create content, it's basically "talking to video" — someone speaking, someone presenting themselves on camera.
Beyond the connection to "video" and "voice," we also have a vision: we hope that in the future, everyone can have their own dedicated space. In this space, you can share your thoughts and emotions through video, just like writing a blog — having your own personal zone. Based on this concept, we named it "Vozo."
But the main reason we chose this name is that we all like how it sounds — short and punchy, vozo.ai is only six letters total, and very catchy. That's how we ultimately settled on it.

Video Models and Vozo: A Differentiated Technical Path
🚥 Koji
The past year has seen enormous changes in the large video model space.
Our last episode was actually a conversation with Luma's product manager about 20 Questions on AI Video Foundation Models, reviewing everything that's happened in the video foundation model space from Sora's launch to now.
Among these developments, which are directly relevant to Vozo, and which are indirectly relevant?
👦🏻 Changyin Zhou
Sora has relatively limited connection to our business.
In October 2023, as I mentioned earlier, we issued a press release showcasing the visual model our team had developed. This was before Runway V2 launched — the period after Runway's first-generation model. Through the practical experience of this project, I formed a clear understanding of the technical bottlenecks that visual foundation models face in video generation.
I could estimate the timeline for breaking through these technical bottlenecks, including key metrics like controllability, consistency, and computational cost. I also assessed the time needed to reduce costs to levels acceptable for ordinary creators — for example, keeping the cost of generating one hour of video within $200–300. Based on these assessments, I decided not to continue developing visual foundation models.**
Of course, other factors influenced this decision as well, such as the substantial capital requirements for this work. I'm not particularly skilled at fundraising, so I felt this probably wasn't something I could do. Therefore, I pivoted toward developing AI-enhanced or AI-assisted video creation tools, rather than text-to-video generation systems. I believed the latter would be difficult to achieve major breakthroughs in within the short term of one to two years.
Additionally, I believed that even if breakthroughs occurred, they wouldn't become industry moats — something that was later validated. While Sora's launch was impressive, far exceeding other products, we anticipated that other companies would release similar technology within months. This indeed happened as expected — now multiple Chinese companies are capable of developing visual foundation models.
This formed a personal judgment of mine:
Whether it's language models, audio models, or visual multimodal models, as long as it's general-purpose, it won't become a moat in the future. Because there will always be open-source and various other pathways to democratize it, so we try to avoid these areas in our entrepreneurship.
All models we develop independently are tailored to specialized needs in specific application scenarios. For example, in translation, we have special requirements for tone preservation, so we specifically developed voice cloning, speech synthesis, and lip-sync models. We continuously iterate product models in vertical domains around real user needs. For external foundation models, as long as they're applicable to our scenarios, we adopt and integrate them.
🚥 Koji
You've certainly done extensive work to improve user experience.
From the original teleprompter app to now "translation tone preservation" and other aspects, this is evident. However, I feel these efforts may not be fully perceived by users?
👦🏻 Changyin Zhou
Yes, I believe users can only truly appreciate the value of these technologies after using them. Take our translation feature, for example — if you try it yourself, you'll discover there are many difficult aspects to translation.
For example, when translating from Chinese to German, the difference in expression length between the two languages is huge. From what I understand, German is probably one of the most verbose languages out there — content that takes 5 seconds to say in Chinese might need 15 seconds in German. In the same video, if the visuals don't change significantly, you get sync issues. If the Chinese portion finishes in 5 seconds, does the mouth keep moving or stop? You can't have the mouth stay shut for 15 seconds, right? How do you solve this?
There are actually many solutions. In translation, we need to find a version that semantically matches the original, tonally approximates the original voice, and naturally matches the lip movements. This essentially becomes an optimization problem.
Different languages have their own characteristics. For instance, when you shoot a one-minute or 15-second short video to tell a brand story, the brand name is usually a proper noun. If you don't understand this during translation, you might incorrectly translate the brand name. With human translation, you can tell the translator in advance: "This is my brand name, please don't get it wrong." But machine translation typically lacks this contextual understanding and will just translate it directly. So we need a reasonable way to guide the translation system to make adjustments, which makes the problem I just mentioned even more complex.
Lip sync presents similar challenges — different languages have different lip movement characteristics. Regarding emotional expression, ordinary voice cloning technology, like if Koji or Ronghui spoke for a minute, might just learn your overall vocal tone from that minute of speech. But translation is different — ideally, we want every sentence's emotion to be accurately replicated. If one sentence in the original is calm and the next is excited, the translated voice should maintain the same emotional shifts.
However, translation can't simply proceed sentence by sentence, as that would lose context and degrade translation quality. So we need to consider context while maintaining correspondence, and also replicate the original voice's emotions. This is why, in industry common knowledge, machine translation has long been considered inadequate.
If you have high quality requirements, traditionally you'd hire a professional team spending $50-100 per minute for translation. But actually, if these technologies are done well enough, I believe the results can surpass average human translation. Of course, compared to top experts, there may still be a gap. But I think in another year or two, this technology might surpass human expert translation levels.
For e-commerce users, when you need to translate a promotional video, you basically just input a video and get a translated version that preserves the original tone, intonation, and emotion. Recently we've also tried translating short dramas, which is more challenging because expressions in short dramas are usually very exaggerated — sometimes characters excitedly slam the table. How to preserve these emotions and tones is a major challenge.
So we're gradually tackling more difficult problems. Initially we started with simple speeches, and now we're slowly able to handle some short drama translations.
🚥 Koji
All the issues you mentioned above — I feel like each one is quite interesting.
When facing these problems, do you prioritize engineering approaches or technical approaches to solve them?
👦🏻 Changyin Zhou
We employ multiple solutions, including R&D approaches (like model improvements), technical approaches (engineering methods), and product approaches. Generally, we prioritize product approaches first — for example, adding a popup to prompt users to click somewhere is usually the most direct and effective solution.
Next is technical-level optimization. Like the optimizations I just mentioned: you need to extend voice duration, align with visuals, and maintain unchanged emotional expression — this is fundamentally an optimization problem. We can write algorithms to achieve these optimizations, which falls under engineering methods.
There are also issues like precise tone replication — how to quickly replicate emotional expression sentence by sentence — which requires iterative model improvements. So solutions fall into these three layers, which makes the work very interesting. When problems arise, we need to decide which method to use, which are temporary solutions for now, and which are improvements that must be completed in the future. The tone replication I mentioned earlier is a good example.
Initially, we gave users some interactive options, like allowing them to intensify certain expressive effects and letting users control it themselves. But this is actually very difficult, especially for translation — many users can't even understand the second language. So we gradually shifted toward using models to directly complete these tasks for users.
There's another interesting problem: for example, translating from Chinese to Arabic — as a user, you might have no idea whether the translation is accurate. What do you do in this situation? If you hire someone to translate, after paying and signing a contract, if they make a mistake, you can hold them accountable. But as a SaaS provider, users can't hold us accountable, so how do we solve this?
Therefore, we provide some innovative features, like "back translation." This feature translates the translated content back into the original language, and then you can compare the original text with the back-translated text. If the meanings are roughly the same, then the original translation is likely accurate.
🚥 Koji
Interesting! First translate into Arabic, then translate Arabic back into Chinese. That's a bit like that game on Happy Camp, where one person blindfolded tells something to another person, and then it gets passed forward.
👦🏻 Changyin Zhou
Otherwise this problem is hard to solve. How do you convince users, especially if they're posting a very important marketing video — it's hard for them to click that button when they don't know if your translation is correct.
🚥 Koji
You just mentioned Sora's release, and it seems visual models don't actually have much impact on what you do at Vozo.
But I feel like over the past year, when people talk about AI video, they all think it's visual models making breakthrough progress, all the news is about them, all the buzzworthy products are about them.
What technological breakthroughs over the past year made Vozo go from impossible to possible, or from only being able to do 60 out of 100 to 80 or 90?
👦🏻 Changyin Zhou
Right, these technologies are all closely related.
Take Sora's DiT architecture, for example — this technology is directly relevant to us. In the voice replication and lip generation field, if you're familiar with this direction, industry insiders all know there was an old technical solution four or five years ago, mainly relying on GAN or other generative models, but the clarity was very low and realism was poor.
After this wave of technological revolution, we started using transformers for lip generation. Recently there have been new technical evolutions, like Gaussian splatting, which can generate higher quality content more quickly. We don't focus on developing very foundational new technologies to replace transformers, but rather optimize lip generation based on existing technologies. After translation, we can also adjust lip movements. Currently, our lip sync technology is probably among the industry-leading, best ones, thanks to our large data accumulation and continuous follow-up with the latest technologies.
We also apply various video generation models. In our latest released features, we can make static images move and speak, which actually uses visual large models for generation. Our model has its unique characteristics, mainly focusing on making photos move naturally. Many companies are researching how to improve generation speed and how to make dynamic effects more harmonious with voice. As the entire video generation industry keeps advancing, we strive to stand on the shoulders of predecessors, follow this development trend, and solve user problems that couldn't be solved in the past.
Returning to your earlier question, whether it's fast single-sentence voice cloning, highly realistic lip and facial expressions, or overall frame generation — most of these technological breakthroughs occurred in the past one and a half to two years, with some innovations only appearing in the past six months.
🚥 Koji
Visual models are developing rapidly — Google recently released Veo 2 to widespread acclaim.
What do you think of this view: as foundation models keep evolving, they may eventually "swallow" and replace products that focus on specific function optimization?
👦🏻 Changyin Zhou
Yes, I think this will definitely happen. It's like a large vehicle. For product developers, we have an internal principle:
If it's a standard model, don't touch it. We should focus on developing things close to the application layer that are distinctive. These differentiated products are actually very stable.
Looking back, take Midjourney as an example — from text-to-image to the overall generation framework, the technical foundation may be largely similar. But from a business perspective, many users have grown accustomed to using Midjourney, and Midjourney itself has many fine technical adjustments that create huge user experience differences from these small variations.
The same applies to video — in the future, there may be more accessible visual models similar to DeepSeek. But when you apply technology to specific scenarios, the differences become extremely significant. This phenomenon was evident in the past Google Glass project, my previous entrepreneurial experience, and including Midjourney's David in his previous venture. In the same technological era, he could do much better than others.
This is precisely the direction that application-layer technical personnel should focus on.
I don't think we need to overly worry about a single model eliminating all technical space and innovation opportunities — that's impossible.
On the basis of similar technologies, there remains huge room for innovation and optimization at the application layer.
🚥 Koji
Are there any ideas where you've seen technological breakthroughs create new product opportunities but haven't had time to pursue them? Perhaps this could give some inspiration to friends who are choosing entrepreneurial directions.
👦🏻 Changyin Zhou
I can't say for sure, because only actual research can reveal the truth. But I'm personally very interested in certain directions. Because I previously participated in Google Glass**, I think combining glasses with low-latency LLMs would be very interesting — it contains huge imaginative space.**
However, this might lead me to repeat my mistakes — getting excited about a technology and diving in. When making real decisions, you still need business analysis. From a technical perspective, many features we wanted to achieve when developing Google Glass were impossible at the time, but are now all possible.
There was one thing in the Google Glass project that I also very much buy into, and that Sergey Brin really wanted to do: "Google Glass makes you smarter." Their vision was: for example, if Ronghui asks me a question, I actually can't answer it, but the glasses can quickly give me prompts so fast that I feel like the answer came from my own mind. This kind of experience is something users like me would be willing to pay for.

From Google X to Entrepreneurship: The Shift from Elite to Down-to-Earth
🚥 Ronghui
You mentioned establishing a lab earlier — I think this is relatively uncommon for startups. Could you share what goals you wanted to achieve at the time, and how this lab helped advance those goals?
Also, I understand you previously came from Google X — was establishing this lab influenced by your experience there? Or could you first introduce what your time at Google X was like?
👦🏻 Changyin Zhou
Although I've been in China recently, most of my career was spent in the United States.
In 2011, as I was nearing the completion of my PhD at Columbia University, I faced the choice of becoming a professor or exploring other paths. At that time, a professor from Stanford University invited me to join a new team he was forming at Google X. So I took a leave of absence from Columbia and built a new team at Google X with this professor and another scholar.
Looking back, our team was essentially created to satisfy the diverse exploratory interests of Google co-founder Sergey Brin. The team eventually expanded to about 12 people, including four Grammy Award winners. We basically assembled the top talent in computing and photography, which enabled us to complete many groundbreaking projects. We were responsible for developing the core imaging and video processing algorithms for Google Glass, laying the entire technical foundation. From a technical perspective, almost all image processing and visual processing technology on Android phones today originates from the technical platform we established at that time, which had a profound impact on my career development.
I then embarked on my entrepreneurial journey. My first startup was in the United States, focused on cutting-edge technology in immersive video — at the time, it was probably the startup generating the highest-resolution video rendering in the industry. My second startup — my current company — presents a stark contrast. This company does very down-to-earth work, which became my biggest lesson:
You must develop features that users clearly need and cannot do without.
These grounded products are often not very "sexy." The first feature I built made me feel somewhat "low-level," even though users needed it. This created a personal emotional conflict that I needed some outlet for. On the other hand, I felt these grounded features, while meeting user needs, couldn't achieve the grander vision of "video expression freedom" that we pursued. Using traditional methods, it would be difficult to reach our goals, even though doing so could generate revenue.
So I realized I needed to establish a research department to solve some core problems. For example, some people have poor on-camera presence, unpleasant voices, or unsmooth delivery — no amount of editing can fix these issues. Even with the best teleprompter and a well-prepared script, they can't produce good content. These problems needed to be solved, so we began our research work.
This decision was somewhat willful, but very fortunate. It wasn't that we achieved breakthroughs alone — the entire industry made multiple major advances in 2023. Our lab leveraged these industry breakthroughs to push related technologies forward. So this may have been a gamble, but also a lucky one.
🚥 Ronghui
This is actually what Steve Jobs said — at some point on the timeline, you'll find that all the dots connect.
What you just said suddenly reminded me — I met you during your first startup, right?
👦🏻 Changyin Zhou
That's right.
🚥 Ronghui
Could you say more — for example, when you worked at Google X, was it an environment with no budget constraints, purely pursuing exploration? Is this an ideal environment for scientific research?
👦🏻 Changyin Zhou
Yes, I don't think you could imagine a better environment.
The working conditions were very generous. I'll give you an example: I was simultaneously managing an image lab, and if I needed to purchase equipment, I could buy it directly for anything under $10,000 — very luxurious.
In terms of hiring, in the first phase we would poach A-plus talent from other Google departments. At the time, Larry Page managed the formal business, while Sergey Brin oversaw Google X, exploring various innovative projects. One day Larry got angry and said we couldn't keep poaching people from other Google departments anymore, so we started recruiting externally. Basically, we would look for the top experts in specific fields, and resources were quite abundant.
However, this also became one of the reasons I later left. Once in this state, we were basically doing research work all the time. Later I led a project with six or seven people collaborating with me. We demoed it at an all-hands meeting, and after the presentation everyone thought it was very cool — and then nothing happened, except it was hung on the wall as an exhibit.
I found this no different from my PhD experience — it felt easy but somewhat wasteful of opportunities, too free. When this freedom went to extremes, I realized the project was difficult to productize and couldn't have real impact. This was the main reason I ultimately chose to leave.
🚥 Ronghui
Could you explain to the audience what Google's A-plus refers to?
👦🏻 Changyin Zhou
We would look for the most capable, highest-performing talent in other departments. For example, if we spotted someone particularly excellent, the smartest member of the Google Earth team (responsible for Google Maps), we would go recruit them. These people were usually willing to join our team, so basically we could pick our own talent within Google.
This practice wasn't actually ideal for the company as a whole, especially for the business departments, because they were responsible for making money, while our side was mainly spending it (laughs).
🚥 Ronghui
Just a couple days ago, I was listening to Marc Andreessen's latest podcast episode where they discussed and highly praised the significance of open source.
They mentioned that it's precisely because of open source that academia gained the ability to do things that only large companies could do until recently — because the costs were too high.
👦🏻 Changyin Zhou
I think putting talented people in places where they can have impact is quite important.
Rather than having some large company or institution gather very talented people together without producing results — I think that's actually quite wasteful.
🚥 Ronghui
At that time, you held the desire to make your research land, to become reality.
Could you talk about what your first entrepreneurial experience mainly involved? Because I remember it was VR, right?
👦🏻 Changyin Zhou
Yes, this is interesting. Although I left Google with the intention of avoiding excessive research orientation, looking back at my first entrepreneurial experience, it was actually still very research-driven. At the time, I thought I was working very hard to grasp user needs, even convinced that I had captured the core needs, but looking back now, that wasn't the case.
In my first startup, I played more of a CTO role, so I was particularly focused on whether our technology was industry-leading. The direction we tried was actually quite adventurous — we wanted to enable two people to meet anytime, anywhere regardless of where they were, a so-called "teleportation" concept. For this, we developed extensive video compression technology, researching how to achieve high-definition, real-time rendering.
From a macro logic perspective, this need seemed enormous — enabling any two people to connect across spatial limitations. But actually, if you analyze it carefully from a business angle, you'll find this business scenario doesn't hold up. There are many factors that would make the business model difficult to realize.
This shows that merely having an idea that seems logically sound and broadly marketable is not sufficient to justify implementing a project.
🚥 Koji
When Apple Vision Pro launched, it demonstrated similar functionality, particularly the FaceTime feature. The demo was indeed impressive — we all tried it, and the initial experience was shocking. But it also seemed to end after just trying it out — everyone's devices started gathering dust, and no one was actually using it for calls. You also mentioned this isn't a real must-have — have you thought about the reasons behind this?
👦🏻 Changyin Zhou
When Vision Pro first launched, I actually wasn't very optimistic about it, even though I knew the experience would be excellent. One employee from my first startup later went to participate in Vision Pro's development because he remained passionate about the VR space.
Whether something can work commercially depends on many conditions. First is the form factor — can ordinary people accept it, and are there alternatives? On the other hand, you need to form a complete ecosystem — you need a large number of apps and content production, the entire industry chain must operate smoothly.
After my first startup, I adopted a more conservative business strategy: focusing on the last missing link in the entire industry chain. Sometimes you might feel a vision is beautiful and should be achievable. But if something requires five links to complete, and you're only responsible for the first link, expecting others to complete the other four — this is actually extremely difficult.
Vision Pro faces similar problems. Whether from product form, price, or must-have reason for users, it's missing many elements. Of course, it does have captivating aspects — the experience is excellent, very cool. You can imagine many beautiful application scenarios, but these scenarios are difficult to form a complete commercial chain, so it's hard to truly succeed.
Even a company as financially powerful as Apple finds it difficult to string together such a massive industry chain. For startups, the best strategy might be to stay away from such grand projects that require building a complete ecosystem.
🚥 Koji
Actually, Vozo's first release was already very successful. Behind this success was user recognition — they enjoyed using it and were willing to spread it.
This also reflects your successful transformation from a researcher role to focusing on discovering and meeting real user needs. When building Vozo, what do you think were the few things you did right that led to this result?
👦🏻 Changyin Zhou
In developing Vozo, I found myself more patient than ever before. Looking back at previous product development, as someone research-oriented, I would often get too excited — once I had a good idea, I couldn't wait to implement it, otherwise I'd feel regret. Vozo's birth, however, went through multiple cycles of conception and rejection.
Vozo's predecessor was a feature I had GPT help me write that I wanted to use myself. This allowed me to do video editing via command line in Terminal on my computer. Around March 2024, I was quite satisfied with this concept and started actually using it to edit and modify videos, only to find the actual results fell short of expectations.
Although this tool could theoretically help me modify anything, I wasn't sure how to operate it. I started asking GPT for modification methods, then adjusted item by item, but this workflow seemed cumbersome. So I integrated GPT directly — I could just tell it "please make the video gentler" and it would complete the modification.
Starting from March 2024, I wrote small programs to use, test, and iterate, continuously adding new features. It wasn't until July that I developed a version I was relatively satisfied with and decided to launch.
Along the way, we conducted extensive research, which also brought benefits. Because we had other products and multiple communities before, I had a good understanding of ordinary creators' skill levels and the problems they might encounter. Although we had accumulated video visual models for quite some time, it still took a considerable amount of time before we actually launched the product.
I think it's worth spending the time to find the right product before promoting it, rather than building something and then realizing you got it wrong.
🚥 Koji
You mentioned using command lines directly in Terminal to edit videos. Was this something you tried during the process of exploring user needs, or did you actually have this practical need at the time?
👦🏻 Changyin Zhou
I conceived an idea — I wanted to edit video the way you edit text. Although I had sketched out the concept, I believed that an idea existing only in your head doesn't count, so I decided to actually build it. The fastest way to implement it was through Terminal operations, so I developed a software without a graphical interface that could execute all editing functions through command lines.
After completing the first video, team members started asking: "Can this video be changed to look like that too?" However, making that video had consumed an enormous amount of my time, requiring frame-by-frame meticulous processing — this was just my initial version.
Then I began thinking about how to reduce production time from 3 hours to around 10 minutes, because I realized that if ordinary users needed more than 10 minutes to complete a task, they would likely give up. So I gradually improved my prototype until, at a certain stage, I felt this project was becoming quite meaningful, and the team truly joined in to develop this product.
🚥 Ronghui
You mentioned your first entrepreneurial experience ended, and later you returned to China to start another business, thinking at the time that you definitely wanted to do something very down-to-earth.
What happened at that time, or what kind of feeling, gave you this idea?
👦🏻 Changyin Zhou
My entrepreneurial experience changed significantly over the past few years. When we developed the VR project, we served major clients including AT&T, Verizon, China Mobile, and the Publicity Department.
But I discovered an obvious problem: After every product iteration, we always had to beg these clients to use the new features. Despite paying for them, many times the products just sat idle, never truly integrated into their workflows. We had to actively ask for feedback, while clients often hadn't used the product at all. This situation limited our previous company's growth potential.
This feeling might not have been so obvious at the time, but when I returned to China in 2020 and stayed longer in 2021, things became clearer. During the pandemic, I was stuck in Hangzhou and used this opportunity to talk with CEOs of more than ten local MCNs. Conversations with these people formed a striking contrast: every time we talked, these creators would raise numerous specific needs, describing in detail how they wanted to make videos.
This formed a strong contrast with the VR project experience: on one side were clients with clear needs that I couldn't yet meet; on the other was a situation where I had developed many features but had to beg clients to use them. The latter experience was especially painful.
This made me realize that a successful business model should provide products people truly crave, that they can immediately put to use once completed — this is what good business experience looks like.
🚥 Ronghui
So what did you do then?
👦🏻 Changyin Zhou
Our initial experience developing the live streaming machine was quite interesting. At the time, many MCN companies wanted to build buildings with hundreds of live streaming rooms, but faced management difficulties. High-end live streaming typically required multiple camera angles and a director, with everyone coordinating via headsets — the entire process was extremely complex.
To solve this, we developed a live streaming machine about the size of a human head. With just one person holding a tablet, most camera switching was automatically completed by the system. It could understand scene content, automatically switching to close-ups of hands when you showed products, greatly simplifying the director's work.
As researchers, we naturally wanted to use AI to replace traditional methods, but while this product was innovative, it still wasn't grounded enough and had various business problems.
Half a year later, we terminated this project and turned to developing what remains a successful product — the teleprompter. Teleprompting is the biggest challenge most people (including myself) face. Once you need to memorize more than half a minute or a minute of content, it's easy to forget. During filming, if you can't remember the content and frequently check prompts, it often ruins the video.
Our AI teleprompter was designed to be simple and practical. It floats above the phone near the camera, similar to karaoke but smarter — the text automatically scrolls with your speaking pace, stopping when you stop and speeding up when you speak faster. This solved a core pain point for many non-professional creators.
This product brought unexpected rewards. Initially I developed it simply because I was certain users needed it, unsure if it could be profitable. But after launch, we discovered the payment rate was surprisingly high. So we gradually expanded around this core feature, adding more functions, and the payment rate continued to increase.
The product launched in 2022 and has accumulated about 8 million users to date. We also built a private community, because many influencers need professional guidance, and now the community has grown to nearly 100,000 people. This made me realize the domestic market is enormous, with very strong demand for grounded products.
Starting from 2021, we first did the live streaming machine, then shifted to short video production, gradually improving the app around the teleprompter. Today, this app has become our main revenue source.
🚥 Koji
Making a teleprompter app at that time — this sounds like something where your research insights from perhaps ten years couldn't be applied.
What was your mindset then? Did you feel a sense of disconnection, as if all your previous professional accumulation suddenly lost its stage?
👦🏻 Changyin Zhou
Indeed, sometimes when I talk with former teachers or classmates, I usually don't mention what I'm doing (laughs), because a teleprompter isn't a particularly "sexy" or high-end product.
But coming back to it, developing a truly good AI teleprompter is actually very challenging. When users record, they may face various complex situations: loud room noise, heavy accents, unstable speaking pace, or jumpy expressions.
To make the product perform well under these conditions requires solving numerous unglamorous but crucial technical problems. Plus the issue of users' widely varying device performance — the entire development process was full of "dirty work."
This situation, as you mentioned earlier, to some extent forced me to later establish a lab.
🚥 Koji
If you could go back to that time, would you still build the lab?
Do you think that if you hadn't looked down on the teleprompter app you developed, and instead invested more time in developing similar products, perhaps today's entrepreneurial progress would be faster and the results more significant?
👦🏻 Changyin Zhou
I think there are indeed multiple possibilities in this situation. That decision was indeed somewhat impulsive. Of course, this relates to my previous mentor, an American foreign academician.
From a logical perspective, we were indeed on the front lines, understanding numerous practical problems in video creation. While many researchers possess strong research capabilities, they often don't know what the real problems are. Therefore, from this macro logic, establishing a deep research lab and determining topics and research directions was valuable. This conclusion itself is reasonable.
The problem might be the timing — perhaps I shouldn't have done this while simultaneously running a startup. At the time, I indeed didn't think it through that thoroughly and just went ahead with it. Looking back now, I'm probably 50% wouldn't do it again, 50% would do it again. I'm not entirely certain if I would establish a lab again if given another chance.
🚥 Ronghui
So in terms of timeline, it was around the teleprompter period.
Following this timeline, from Google X to the first VR startup, to Hangzhou for the live streaming machine, then the teleprompter, then the Research Lab, and finally Vozo merging with the teleprompter app.
👦🏻 Changyin Zhou
Right, the teleprompter has actually become a feature within the original app. Because it was just the initial entry point of this app.
But going back to Koji's question just now, I think if I went back to that time, I would most likely still build the lab. If I didn't do the lab, I would definitely do something else.
🚥 Koji
Something else more crazy?
👦🏻 Changyin Zhou
Right, otherwise I think if I were just doing purely grounded things that can make money, I probably wouldn't accept that.
🚥 Koji
What about now? Do you feel you've accepted it now?
👦🏻 Changyin Zhou
I think Vozo's performance has at least given me a sense of pride — it's something I created with my own hands, and it makes me feel satisfied. In comparison, if I had only developed the teleprompter, I might feel I couldn't account to myself for this entrepreneurial journey.
🚥 Ronghui
I think doing the teleprompter at that time was very demanding for your career experience.
Because you had to change past work habits. As a returnee with a halo, the people you contacted after returning weren't necessarily doing grounded things themselves — you had to engage with a group of people you might never have contacted before, people you previously wouldn't have known how to interact with.
My question is: during this period, what kind of self-reflection or other significant adjustments did you make that had a relatively big impact on you, enabling you to do things you'd never done before, overcoming the fear that comes from never having done something?
👦🏻 Changyin Zhou
I think my personality is quite interesting — because I hadn't done this before, I actually found it quite exciting when doing it.
Sometimes going to live streaming bases and talking with people I'd never talked to before would really surprise me. When I first started, for example, a user complained that our teleprompter wasn't working well. We asked if his room was noisy, if the image looked poor, if the lighting was dim. He would confidently tell me his room was very quiet and the lighting was excellent.
We were puzzled, thinking we had a bug, so we went to his filming location and found the lighting was very dim, with cars passing by noisily nearby. I found it very interesting — he wasn't lying, that's just what he believed. People are very different. He thought the room was bright, but the bright we meant wasn't the bright he meant; we thought quiet, but not the quiet he meant. This is fascinating.
So when I went to many live streaming bases, including talking with every MCN CEO, I felt they were completely different from us, and it was fun. The fun was one aspect, but sometimes when I quieted down at night and thought about what I was doing, I would have other feelings. I know some people might find this hard to accept, but I was okay — this part made me feel exciting.
The unexciting part was that what I was doing seemed like others could do too, or maybe I could do it slightly better, but others could do it if they tried. This was my challenge. Because researchers or scientists generally have this idea that I need to do things others can't do. This was a relatively big psychological challenge.
🚥 Koji
This is also an economic rationality consideration.
When what I'm building is something ten thousand other people could build, I don't have any unique competitive edge. So I need to do what others can't do — that kind of competitive advantage gives you sustained differentiation and makes things easier over time.
👦🏻 Changyin Zhou
Right, for people coming from a technical background, this hurdle is especially hard to get past. You always feel like if what you're building doesn't have a technical lead, then it's not worth doing.
Sometimes we can't really call ourselves elites, but if we're talking about elite entrepreneurship, this is actually a very difficult thing to break through. You always feel like you need to do something different, but from a business perspective, that's not really how it works.
🚥 Ronghui
I think this is something many people with research or technical backgrounds encounter when they start companies. How do you look at this from a business perspective rather than a technical breakthrough perspective?
👦🏻 Changyin Zhou
Right, I see projects like this every day. I have a few thoughts, though they may not be particularly systematic. First, you need to become a good product manager — you have to abandon your wishful thinking.
For example, my first experience was more like wishful thinking. I thought if I built a system for remote transmission, people would use it, and then someone would build cameras and equipment for it, and everyone would pay. These ideas seem logically sound, but they don't actually happen — or whether they would happen, you could just ask. That's the first thing to overcome.
Second is knowledge — entrepreneurs often don't know what percentage of the market actually needs your innovation. If you really do the research, you'll be very surprised: those innovation points you care so much about might only matter to 1% of users. This is a knowledge gap.
So one is an attitude problem of wishful thinking, one is understanding the market better, and then there's how to shed your ego — this needs a systematic theory. I don't have such a theory right now. Maybe Koji, you could try to summarize one — I think it would be very helpful for many entrepreneurs.

Shedding Ego: The Entrepreneur's Self-Renewal
🚥 Ronghui
You just mentioned shedding your ego — I think this is the hardest part. Looking back now, what did you actually do to shed your ego?
👦🏻 Changyin Zhou
It was mostly passive lessons that made me do it. Because you don't think you're wrong, but after being wrong a few times, you figure it out.
🚥 Ronghui
Were there any particular moments when you felt you were going through a major change?
👦🏻 Changyin Zhou
You mean a specific point in time?
🚥 Ronghui
Or some special experience, or at this stage you required yourself to do things you'd never tried before.
👦🏻 Changyin Zhou
I think at least behaviorally there were some changes.
One manifestation of ego is thinking that whatever you think is right — whether big or small, you'd try to convince others. I'm not sure when this change happened, but in the team, since I'm still quite involved in product and technology, sometimes I'll propose a technical solution and the younger folks might reject it.
Now I'm generally quite used to being rejected. Even if their rejection isn't necessarily right, as long as it's not very critical, I'll let it pass. This is a change — I wasn't like this before. I used to think I was the smartest and my ideas must be right. And I used to think these details mattered a lot — if we did it that way, performance would drop from 99% to 98.9%, and that was unacceptable (laughs). But I can't really recall when this change started.
🚥 Koji
Was it because letting go like this gave you positive feedback at some point?
👦🏻 Changyin Zhou
I think after letting go, I had much more time.
Because probabilistically speaking, if we used my solution it might be 70 points, and with their solution it might be 65 points — there's not much difference. And because it's their solution, they'll execute it better, and the result might even be better than my solution. So there's no need to obsess over this kind of thing.
Only for some truly critical things do I need to think very clearly and convince everyone, and those should be very rare.
🚥 Ronghui
Did you have any new understanding of entrepreneurship at this point?
👦🏻 Changyin Zhou
I had a mental journey when I first started my company. I left Google in 2015 to start my first company. At that time I was quite naive — I was the CTO and just solved technical problems. So entrepreneurship was this vague thing, and I just went and did it because it felt exciting.
Later I gradually felt like there were so many things in entrepreneurship, busy every day. Including when I was CEO in my second startup, I would do everything myself. But my energy was actually very scattered, and I don't think I made the right decisions on some important company matters, probably because I didn't spend enough energy on them. Then I gradually realized there are actually only a few important things — now I struggle more with figuring out which things are the important ones.
For example, now I have three important things, but deep down I know maybe only two of them are truly important, and I'll spend a lot of time thinking about which is more important. So I imagine some more impressive entrepreneurs can tell at a glance that this thing is more important and that thing doesn't need to be done.
I don't know how this path will evolve over the next three to five years, but I think focus — knowing what's more important — might be the difference between those particularly impressive entrepreneurs and more ordinary people like me.
🚥 Koji
I'm curious — when this company was fundraising, you must have talked to a lot of people like Linear Capital, Sequoia, the whole circle. Did you raise money with the live streaming device idea?
👦🏻 Changyin Zhou
Right, the live streaming device.
🚥 Koji
Your entire research background, including your first startup which was doing VR, and then you went to do live streaming devices — this was indeed a huge shift.
How did you make up your mind at the time? Did you see some grand commercial prospect in "live streaming devices" — something that made you willing to step up as CEO and take the biggest risk to start a company?
👦🏻 Changyin Zhou
Actually there were two lines of thinking. First, I felt that domestic e-commerce and demand for short video — this is a very large market, and there must be opportunities in it. People with some technical capability, plus business capability (since I'm from Wenzhou, I always feel my business sense can't be too bad), I felt there must be opportunities here.
Whether the live streaming device was the optimal choice, I wasn't sure, but at the time it seemed like there were clear customers who wanted it, and we could sell it for quite a bit of money. It's just that at the time, I didn't fully analyze the long-term development path and potential obstacles for this type of hardware-software project before jumping straight into action.
So in the middle we had one fundraising round where a fairly well-known head of a domestic dollar fund directly asked me: "Why are you doing this? Can you do something else?"
🚥 Koji
I think this was indeed quite a surprising life turn.
Doing research for so long, and then when you come out to start a company and raise money, you tell this kind of story. Everyone would suspect they heard your direction wrong — there would be this huge sense of dissonance.
Did you often do this as a child, making decisions that surprised people?
👦🏻 Changyin Zhou
Right, I'm a bit unusual — I've made many strange decisions.
🚥 Koji
Have there been other "outrageous" decisions in your life?
👦🏻 Changyin Zhou
My undergrad was in management, so I studied in the business school first, then after graduation went to work at Microsoft. This was a very strange thing. After working at Microsoft for a while, I felt I wanted to do research, so I resigned from Microsoft and went to grad school for my master's and PhD.
🚥 Ronghui
Undergrad was management, then grad school was computer science, right?
👦🏻 Changyin Zhou
Right.
🚥 Koji
And most people at that age switching to research — it's actually very, very difficult.
👦🏻 Changyin Zhou
I don't worry too much about these things — maybe sunk costs aren't very important to me. I feel like if it's what needs to be done next, then you can just do it.
🚥 Ronghui
You're that Li Dan quote: "Sunk costs don't participate in major decisions."
I think what you said earlier about the things before Vozo — because you were doing some research, then using tools to write it out yourself. Actually the previous research might have had some environmental advantages, and you might have been relatively less exposed to more grounded things earlier on. Then later it became a closed loop, which actually happened to be your own research habit.
Combined with tools, especially the major developments in opportunities and tools related to AI, it all came together to play a role.
👦🏻 Changyin Zhou
Right, I think the final convergence point was actually on product. I think product manager is really quite a difficult role to do well — over all these years, I basically forced myself to become a product manager.
👦🏻 Changyin Zhou
I think product manager might be a quite interesting role in this era.
You need to understand technology, understand the market, even understand how traffic comes in. And when you can combine all these things well into one thing — that's product. So people from technical backgrounds coming to do product, or people from marketing doing product, will face many challenges.
The path I took was probably from research to technology to product — I think it's pretty good, quite an interesting process.
🚥 Ronghui
What was the goal when doing these things? Was it to build a certain kind of company, or to make money?
What was the core psychological motivation? And that "sunk costs don't participate in major decisions" you mentioned — I think that's quite a special quality.
👦🏻 Changyin Zhou
I think that might just be personality — I'm probably a purely rational person.
I'm a probabilist, so it's fine. I think the original intention probably had two parts. First, from the influence of doing research before and being at Google — being a researcher, hoping that my own intelligence could very positively influence many people, influence the world. This was probably the big inner idea.
On the other side, more specifically, when I was at Google I always felt that using video to transmit information was an inevitable thing. Because video has the most information, the highest bandwidth — this would happen sooner or later. I always felt this would definitely happen, and hoped to be one of the main people making it happen.
But in 2015 it was way too early — the market wasn't ready, the technology wasn't ready. By 2021, I felt like there was a bit of an opening. So circling back to Koji's earlier question about why I returned to China in 2021 to do this: this whole video storytelling thing connects back to what I originally wanted to do.
🚥 Ronghui
To sum it up, it's because there's something you deeply believe is bound to happen, and you want to be part of it — ideally, one of the people pushing it forward.
🚥 Koji
You've worked with some of the smartest people out there, seen plenty of top-tier talent. What do you think separates the truly exceptional from everyone else?
👦🏻 Changyin Zhou
I think I've been relatively fortunate to have met some particularly high-profile people. I started at Microsoft Research Asia — not sure if I should name names on the show — but he was my advisor, now a foreign member of the US National Academy of Sciences. I had close interactions with him and observed how he operated. Later he sent me to Microsoft's US headquarters, where I talked with several of the key people there. Then I went to Columbia to work with another academician, essentially the top professor in computational imaging. After that I joined Google, where I worked with Sergey Brin and another Graphics Fellow.
I think they share certain traits. They're extremely focused, and they think about surprisingly few things.
Take my advisor — he takes very few students. He's approaching 70 now, and this year he still published two best papers. He's extremely focused. He identifies the most important problem in his field, then the most important sub-problem within that. He just works on that. Once solved, everything else naturally falls into place. Because once you've solved the most important thing, resources and people naturally gravitate toward you, and the thing gets done. Sometimes it seems almost effortless — just this intense focus. I think that's one defining characteristic.
Many people who aren't at the top may not have the luxury of only doing important things. Life forces them to juggle many other things, which becomes a cycle. Top people only do the single most important thing; they let others handle everything else, or simply don't do it. I think that's a major difference.
But enabling this focus requires capabilities too. For instance, you might want to focus but can't figure out where to focus. Even if someone gave you a million dollars with no strings attached, you might still not know what your most important thing is.
I consider this a fairly important differentiator — possibly one of the most critical. It's something I've been thinking about a lot recently. My views may evolve, but for now I believe this matters tremendously. There's a common psychological tendency at play: people gravitate toward moderation. When considering three different things, our instinct is to assume they're all roughly equally important.
But if you believe 1 is more important than 2, and 2 more than 3 — say you rate them 80, 60, and 40 — I think you're not spreading the variance wide enough. You're probably underestimating the gaps. If you say 80, 60, 40, the reality is more likely 90, 20, 10.
People tend toward the middle.
🚥 Ronghui
What methods do you use now to distinguish what's most important?
👦🏻 Changyin Zhou
One thing I ask myself is: what happens if I don't do this? Often, nothing much happens.
Not "I feel uncomfortable not doing it" — but will our revenue actually drop? Will users actually leave? How many — two users or 20%? Rough it out, and often it turns out to be unimportant.

Learning from the Best: Growth and Mindset
🚥 Ronghui
What methods do you use to keep learning?
👦🏻 Changyin Zhou
These days I mainly learn from GPT. I'm a devoted ChatGPT user. They've probably lost a lot of money on me — I use it every day (laughs).
When o1 first came out, I'd basically burn through my quota every few days and have to wait until the next day. Now I can use it freely. I think it's already smarter than humans, so just learn from it. That's one approach.
The other is to find the strongest person in any given field and learn from them — in academia, just reach out and talk. I think this is quite important: whatever you're doing, find the best person at it and have a conversation first. That's a pretty effective method. Maybe this traces back to my business school days, when I skipped so many classes — and even when I did go, I'd first find the professor to have them highlight the key points for me (laughs).
🚥 Koji
This is a fascinating perspective. Our previous guest Justin, who founded Moonton and later sold it to ByteDance for over $4 billion, gave a similar answer when we asked him — that "learning from the best people" was key.
When we asked him who he planned to learn from next, he mentioned he'd already scheduled a meeting with a partner at DeepSeek for the next day.
But here's my question: not everyone can easily access top talent. How should young people find and approach those they consider impressive?
👦🏻 Changyin Zhou
I think just seeking out the most impressive person within your reach gets you 80% of the way there — you don't necessarily need the absolute best in the field. You'll find many people are actually quite accessible; if you reach out, they'll probably talk to you.
I actually realized this quite late myself — only gradually during grad school. I was at Fudan for my master's, wanting to do computer vision research, and wondered where to go. I wanted to study abroad but didn't know how. So I looked around and saw Microsoft Research Asia in Beijing. I sent an email to a researcher there.
He became an important mentor figure. He called to interview me, and after the call I went to Beijing. There, he introduced me to the senior leader at Microsoft Research Asia I mentioned earlier, who then recommended me for a PhD at Columbia, and later to Microsoft. From there I started attending conferences and giving talks.
Here's a funny story: I gave a research talk, and in the open Q&A an old man asked me something. I answered, and that old man turned out to be my future boss at Google X. He remembered me and later called to ask if I wanted to join him. So I think just paying attention to people within your reach, getting to know them — the network is actually quite small. That's enough.
🚥 Koji
I've heard this before: Treat every conversation like an interview.
But thinking about it more, just relax a bit going into these conversations, try not to be too shy, express yourself more.
👦🏻 Changyin Zhou
Yes, I think this is a killer skill that should be taught in college or even high school.
Why is this worth discussing? Because when we hire domestically in China, I notice our Chinese colleagues are noticeably weaker in this awareness compared to our US colleagues. So I sometimes spend energy thinking about how to help some of the particularly talented ones among them become even better.

Looking Ahead: Domestic Market and Team Expansion
🚥 Ronghui
We know Vozo initially launched on the overseas App Store. Do you now have plans for a Chinese version?
👦🏻 Changyin Zhou
We've actually discussed internationalization strategy for a long time internally — debating whether to support the domestic market and when.
Though we never officially released a Chinese version, we've accumulated a substantial Chinese user base. This probably stems from our position within China's tech network, plus the large base of Chinese short-drama exporters and e-commerce exporters. So many Chinese users have been drawn to our product.
These users fall into two broad categories: some keep requesting a Chinese version; others face payment difficulties because our payment system doesn't support Alipay or WeChat Pay, preventing them from completing purchases. These two issues constitute the bulk of user feedback.
I think the timing is about right. We've more or less completed our PMF iteration, and now we're moving into growth. For the domestic market, I think we should support it. Another debate some companies have is whether to exclude China entirely — our team has never thought that way. We just debated prioritization: China versus Japan versus France, and in what order. We've now decided to support the domestic market regardless, so domestic users can at least understand the interface, pay, and submit support tickets. That's important.
If you're interested in AI video — whether in product, R&D, or engineering — feel free to reach out anytime. We can create roles around the right people.
🚥 Ronghui
Well, thank you so much today, Changyin, for sharing your journey building Vozo, your views on the industry, and so much of your personal experience — especially your transition from researcher to founder, and the reflections and thinking along the way.
Let's wrap here for today. Thanks for joining us at Crossing, and we hope to have more conversations like this in the future 👋.
👦🏻 Changyin Zhou
Thank you Ronghui, thank you Koji — really enjoyed today's conversation ❤️.
🚥 Koji
Thank you, bye bye 👋.

Subscribe to the Crossing Podcast
🚦 We track how the new wave of AI technology is reshaping industries and creating fresh entrepreneurial opportunities. "Crossing" comes from Steve Jobs' metaphor for Apple — standing at the intersection of technology and liberal arts, where great products are born. As AI transforms every sector, we seek out, interview, and bring together the "active doers" of the AI era. Together with them, we explore and embrace what's new and what's possible.
👦🏻 Host Koji: Co-founder of The Fair and Tangdao. I believe technology, especially AI, will fundamentally reshape society and empower humanity. I'm always open to conversations, idea exchanges, and connecting the next possibility. Koji on Jike[3], Koji's website[4]
👧🏻 Host Ronghui: Works at a tech VC, former Silicon Valley correspondent for CBNweekly. Ronghui on Jike[5]
References
[1] Luma.ai: http://luma.ai/
[2] Vozo: https://www.vozo.ai/
[3] Koji on Jike: https://okjk.co/0JSUes
[4] Koji's website: https://koji.super.site/
[5] Ronghui on Jike: https://okjk.co/0cbnYV