Code Moment | Sand.ai Closes New ~$50 Million Funding Round, Doubling Down on Product + Model Dual-Track Strategy
Audio-driven video, dual-engine monetization reaching eight-figure ARR.


MaHui member Sand.ai recently closed a new funding round of roughly $50 million. Its core product VidMuse, since launching earlier this year, has surpassed $10 million in ARR (Annual Recurring Revenue), making it the fastest product in the Video Agent track to reach this commercial milestone.
Recently, Sand.ai founder Yue Cao and VidMuse product lead Zake Zhang Zihe sat down with GeekPark for an in-depth conversation, sharing their latest thinking on the Video Agent path — how Sand.ai uses "music" as its anchor to simultaneously steer both product and model development, moving toward a long-term vision of a "digital production team."
Below is the full GeekPark interview:
In 2026, even as Agentic AI — represented by OpenClaw — has become the "super consensus" across the AI world, video models have begun moving in a different direction.
In the United States, OpenAI has shut down Sora as a standalone product, clearly deprioritizing video generation in its current roadmap. Startups like Runway and Luma AI have also shifted their narrative center toward "world models."
Domestically, the picture looks quite different: video models are becoming a multimodal capability that major tech companies see as essential for their next phase. Whether it's ByteDance's Dreamina or Kuaishou's Keling AI, these video generation products are penetrating deeper into professional creator communities with stronger willingness to pay, beyond mass entertainment scenarios.
Sand.ai is a noteworthy startup sample amid this divergence. Their core product VidMuse features a "Music in, Video Out" product form, placing audio at the most central input position. According to available information, VidMuse has surpassed $10 million in ARR since its launch earlier this year.
Recently, Sand.ai announced a new funding round of approximately $50 million. GeekPark also met in person with founder Yue Cao and VidMuse product lead Zake Zhang Zihe. In Sand.ai's view, music's importance doesn't stem from its correspondence to a particular content type or user base, but from its potential to become a more foundational input starting point for AI-era video creation — one that naturally connects to stronger creative intent.
Meanwhile, Sand.ai has firmly chosen a "dual-wheel drive" path of building both product and model: first using the best-performing models on the market to find product-market fit for their product, then bringing their own models back to critical junctures to swap in better effects, lower costs, and improved margins. From the perspectives of energy, capability, and resources, this is not an easy path. But in Cao's view, this is precisely an advantage startups have over giants: here, model and product can more easily serve the same goal without splitting apart from each other.
And what this path truly points toward is not just a more powerful video generation tool, but a "digital production team" capable of long-term collaboration. Under this new Video Agent product form, the user becomes more like a "financier": no longer needing to act as director and repeatedly prompt for scenes, but able to entrust their creative goals with confidence to a trusted, continuously callable creative partner.
Below is the edited interview:
01 The "China-US Divergence" in Video Model Generation
GeekPark: Have you noticed that HappyHorse thing that's been going viral lately?
Yue Cao: I saw it, pretty interesting. A lot of people asked whether that was our model based on some analysis post on Twitter. Later I found out that some website had directly converted our Magihuman tech report (Sand.ai's latest open-source model) into a webpage, and the name was HappyHorse. (Laughs) But our new model is in training, we'll release it as soon as possible, and there's a high probability we'll open-source it directly, hoping to accelerate the entire industry together.
GeekPark: So it was fake news. But recently you've been internally testing the new product VidMuse 2.0 while also open-sourcing a foundation model — from the outside, this looks like a somewhat contrarian decision. Today everyone is emphasizing commercialization and closed-source; why did you choose open source?
Yue Cao: I think one of the essences of open source is to enhance brand value, and sometimes it can also lower customer acquisition costs. For example, with the DeepSeek-R1 open source, people may not have initially expected it to have such great effects and play such a good role.
For us, when we released Magi-1 last April, we open-sourced that model, probably one of the earliest teams exploring world models. Magi-1 was an autoregressive video foundation model. Zake was still studying in Northern Europe at the time, and he found us after seeing this open-source model.
GeekPark: Today many investment institutions also find entrepreneurs through open-source projects on GitHub. So what stage has the video model track reached today?
Yue Cao: This has entered a stage of "rhythm divergence": some directions will mature first, others will mature later. What's most clearly established now is using video models to replace live-action filming.
In the past, if you wanted to produce a piece of content, you needed to rent locations, lighting, actors, then go through the filming process; now it's increasingly becoming "write the prompt, click generate." This capability first serves professional creators who were already producing content, helping them replace the live-action filming they used to do.
Therefore, what's most mature at this stage isn't general entertainment consumption, but content production with clear goals. As model capabilities strengthen, the proportion of these creators using AI continues to increase, and these people already had production needs and are more willing to pay. Over the past nearly two years, the growth of Keling AI, Runway, and Seedance has all been built on these kinds of scenarios, with typical applications including short video content, advertising and e-commerce, short dramas, and other general content production.
GeekPark: What are the overall differences in how China and the US approach video models?
Yue Cao: I believe the differences between Chinese and American teams essentially stem from different industry and product environments over the past decade.
In North America, over the past decade, the big consumer money has largely been captured by giants like Meta, with relatively fewer startups truly centered on consumer products, so a large number of startups have been more accustomed to making money through B2B SaaS.
Over the past decade in China, product forms like WeChat and short video have been the hottest products, so the overall market has stronger intuition for consumer scenarios. Therefore, on the matter of video generation, Chinese companies place more value on it and believe more strongly that it can quickly generate commercial returns.
To some extent, I understand OpenAI shutting down Sora as shifting more compute resources toward the coding direction. By contrast, Chinese companies appear to value video generation more, because it's already one of the clearest large-scale scenarios beyond coding, and its commercial value is easier to verify.
Sand.ai founder Yue Cao
GeekPark: Specifically, what have large companies and entrepreneurs done? Have you been following Runway's recent moves in the US?
Yue Cao: We actually haven't been paying particularly close attention to Runway. Because it seems that at the product level of "pure video generation for creators," they don't appear to be making particularly large-scale investments anymore, and their overall narrative is increasingly tilting toward "world models" — Luma AI is the same way too. Rather than continuing to strengthen products, American entrepreneurs are more focused on strengthening models and the direction of continued model evolution.
GeekPark: So they're "weakening product, strengthening model"?
Yue Cao: Yes, I think that's the trend in Silicon Valley.
In China, products will enter the commercialization phase more quickly. Represented by Seedance and Keling AI, Chinese video models can achieve paid closed loops faster. However, although there's still a gap between domestic and international leading levels in language models, in the video direction, I believe Chinese companies' model capabilities are already in the world's first tier, which is also why they can more easily pioneer commercial scenarios.
02 One of the Few Technical Consensuses: Audio-Video Co-Generation, Multi-Shot Narrative
GeekPark: Have technical approaches in video models converged today?
Yue Cao: They haven't converged. At least there hasn't yet emerged a unified direction like coding in language models, where everyone must obsess over it and can't afford to fall behind.
The current competition in video models is more like different teams making strengthened choices in different directions. For example, on multi-shot narrative, Seedance is currently in a leading position, but we believe this doesn't come from an unreplicable absolute technical moat — it's more a matter of "choosing this direction earlier and executing it better sooner," thereby gaining a lead cycle of roughly three months.
Actually, looking at model capability progress over the past two to three years, capabilities developed by one company are often followed by others in a very short time, as fast as two to three months, or as slow as three to six months. So the core of competition isn't entirely long-term technical moats, but also includes phased judgment and choices.
GeekPark: Then what have been the most critical technical breakthroughs at the video model level over the past year?
Yue Cao: I believe it's audio-video co-generation and multi-shot narrative.
Google Veo 3 was one of the earliest models to achieve audio-video co-generation, and we followed up quickly too. Its key value is that basic human performance becomes more nuanced and realistic, especially the synchronization between lip movements, voice, and actions, making people look less like AI-generated figures and more like real performances.
GeekPark: What about multi-shot narrative?
Yue Cao: The importance of multi-shot narrative is actually something the industry only suddenly realized after it was achieved. Because it significantly improves the quality and realism of narrative video.
If it's just single-shot generation, even if the image itself is good, people will still vaguely sense that "something's off." Because humans naturally live in 3D space and have very acute perception of whether space is real. Multi-shot narrative allows the same scene to be shown from different perspectives within a short video. For example, first filming one person speaking from one angle, then cutting to another angle to film another person responding. This way, viewers quickly establish a sense of space for this scene, and the overall result feels more real and comfortable.
Additionally, there is a great deal of naturally aligned information in the real world. Image and sound are aligned; different perspectives within the same space are also aligned. In the past, if models only processed single-shot, soundless content, it was equivalent to not utilizing this naturally existing information in reality. Once these different dimensions of information are fed together into the same model, generation quality improves significantly.
GeekPark: This sounds like a process of continuous dimensional elevation — from static images, to dynamic images plus sound, to multi-perspective expression within the same space, with capabilities stacking layer by layer. Stack to a certain critical point, and users suddenly feel "this thing is actually usable."
Yue Cao: This is essentially the nature of multimodality: taking information that was already aligned in the physical world and unifying it through the same model.
GeekPark: In the video model field, will there emerge something like coding is to language models — a "crown jewel"? If so, what would it be?
Yue Cao: If you ask me to give a fully converged answer right now, I don't think there is one yet. But I believe a very critical direction for video models next is likely stronger contextual understanding, thinking, and the more nuanced performance capabilities that result from this.
Today's models can already do some things. For example, you give it a photo and a relatively specific description, and it can already make this person speak a line with a certain emotion, with image and sound generated together, so the alignment is relatively high and it feels quite real.
But this is still at a relatively coarse level. If you want to express audio-visual content more delicately, I think what models need isn't a simpler one-to-one mapping, but thinking. That is, after seeing an entire prompt, it doesn't directly map "say angrily" to an expression, but first understands the context: who is this character, what happened before, what is this scene, how should they express themselves. Only this way will the performance be more nuanced and more fitting to the scene.
No model can truly do this yet, but I think it will come very quickly, and it will be the next very critical breakthrough.
03 From Creator to "Video Investor"
GeekPark: Let's talk about your new product VidMuse 2.0 that's in internal testing. I read your introduction — the interaction logic is "Music in, Video Out." What's the core upgrade this time?
Zake Zhang Zihe: The core of VidMuse 2.0 isn't adding a few more features, but rebuilding the agent framework.
Previously, many Video Agents on the market, including our own 1.0 state, were more like agents "wearing shackles": they could only follow your preset workflow, step by step.
But video creation itself isn't a linear process; it's a very divergent one. So the core upgrade of 2.0 is shifting from this workflow-style, heavily orchestrated tool toward a more open Video Agent. What we want to do is release as much as possible those handcuffs and shackles placed on AI, letting it exercise its own intelligence, flowing along with user needs and the creative process.
GeekPark: Now everyone is starting to loosen the reins, reduce orchestration, and leave more to the agent to create a good environment. VidMuse 2.0 is basically moving in this direction, right?
Zake Zhang Zihe: Yes, because video creation itself is very community-driven. New playstyles, new creative habits, and new forms of expression constantly emerge in the community. If every time a new idea bubbles up in the community, I have to rely on human and material resources to iterate a new feature, this product will never keep up. Even with various coding agents improving efficiency, you can't really be online 24/7 to manually support all these changes.
So from the product perspective, tying AI to fixed workflows can't keep up with the speed of creative evolution.
GeekPark: Since you consider it a Video Agent, what is its benchmark?
Zake Zhang Zihe: From the start, we never treated it as a single-point tool, but as a "complete contractor" or "production team." We see many AI-era creators, in order to make a complete video, shuttling back and forth between DeepSeek, Midjourney, image generation tools, and video generation tools, building their own pipelines — the barrier is very high. The opportunity we saw was: can we build an agent on top of these tools and turn it into a complete production team? Users no longer need to shuttle between various tools themselves; they just state the goal, and the agent organizes the workflow, dispatches agents, and finally delivers the video.
GeekPark: Under this form, the user actually becomes a producer or investor. "Burn" tokens, then get a satisfactory finished film.
Zake Zhang Zihe: Yes.
VidMuse product lead Zake Zhang Zihe
Music as the Starting Point for AI-Era Video Creation
GeekPark: I've heard some people view VidMuse as a vertical product for the MV scenario? You must be aiming for a general-purpose goal, right?
Zake Zhang Zihe: I want to specifically clarify this. We have never internally said we only do MVs, nor have we ever positioned ourselves as an MV Video Agent.
At first we also took some detours. The initial thought was: model capabilities are general-purpose, so the product should also be as general-purpose as possible, without too many presets for the model. But when you actually do it, you find that if you try to cover all scenarios, the product struggles to cross that threshold where "users are willing to pay," so it must converge.
The question is how to converge. Many people slice by content type: music, manga dramas, advertising, making different products respectively. But I don't quite agree with this approach. Because if you box the product in by content type, when it later needs to radiate to more scenarios, it often requires reconstruction. What we ultimately chose wasn't slicing by content type, but slicing by creative chain. That is, I don't first define "I make MVs," but first define: in AI-era video creation, along what chain does it actually proceed.
GeekPark: So you follow "creative intent" to find users? Why does music become a better entry point?
Zake Zhang Zihe: I'm increasingly feeling that audio is a more suitable continuous information entry point than images and text. Images and text are more discrete, but audio, especially music, flows continuously.
We've scrolled through many purely AI-generated videos that went viral on Twitter and YouTube, and found they have a very obvious commonality: many works actually use music or audio to drive the entire creative chain. So that's when I said music is actually like the skeleton of this video.
So I would think: AI-era video may no longer need traditional CapCut-style software logic, but is more likely to proceed along an audio-driven chain. We later chose to start from music not because "the MV category itself," but because I feel that within audio, music occupies a very large part — it's the most natural entry point.
GeekPark: If we extend outward from this logic? What else beyond MVs could there be?
Zake Zhang Zihe: This understanding later extended to advertising as well. I feel that in advertising, many truly memorable things aren't just images and copy, but also melody. A word paired with an earworm melody, paired with simple but strongly memorable images — information transmission gets significantly amplified.
GeekPark: So from a longer-term perspective, you would view "text, image, melody" as a higher-dimensional content format, rather than treating music as merely a subordinate element.
Zake Zhang Zihe: Yes.
VidMuse product interface
GeekPark: Does choosing "Music in" relate to user persona?
Zake Zhang Zihe: Yes, and it's very related.
We have a very clear judgment: many Video Agents hit growth bottlenecks because it's very hard to create users' "creative intent" out of thin air. If a person didn't already have the willingness to produce video, it's hard to suddenly make them start doing this, and ROI is hard to justify positively. But starting from music is different. People who have music naturally already have creative intent; getting them to transition smoothly from music to video, the ROI for acquisition and growth is more positive, which is also one reason we've grown relatively fast.
So music wasn't chosen randomly as a traffic entry point, but directly relates to "creative intent."
GeekPark: What does your current user profile look like?
Zake Zhang Zihe: I'd roughly divide them into two categories.
The first is music-related users, whether traditional musicians or AI musicians. The latter actually makes up a very large portion — for example, Suno gave them creative capabilities, and they grew from being mere music enthusiasts to frequently releasing their own songs, hoping more people will hear them.
But having music alone isn't enough. If you post music on Spotify or SoundCloud, the number of people who actually hear it is still limited; places with more traffic are TikTok, Instagram, YouTube. So they naturally need a video medium. So the first batch of core users I saw were actually: they're very good at making music, but can't make music videos. They were already very professional in the music modality, and came to VidMuse to fill in that step "from music to video."
GeekPark: What about the other category?
Zake Zhang Zihe: We internally call them people doing general lifestyle creation.
This group's creative content leans more toward life and personal expression, like year-end party videos, children's growth, friends' birthdays, family anniversaries — these all count. This direction itself was a new discovery, because this group was actually easily overlooked in the past.
What impressed us even more was that some within this group have very strong personal emotional expression. Some use it to create videos about childhood, family relationships, and other themes. They often already have a song of their own, then use this product to match that song to the scenes they truly want in their hearts, adjusting again and again. Some of this content isn't even posted to any platform — it's not for distribution, but for expression and release.
One important thing about this user group is: they often upload very private photos and stories. They may not be willing to hand this content to a human creator, but are willing to entrust it to a tool or agent to complete. So I feel this is no longer just ordinary content production; it's closer to personal commemoration, emotional processing, even a kind of self-healing creation.
Seeking Certainty
GeekPark: If through orchestration and adding skills, one could use OpenClaw to make a similar product, then what role does your own model play in VidMuse? Is the relationship between your model and product strong coupling or weak coupling?
Yue Cao: Internally, we've been dual-wheel drive from the start.
Product shouldn't be constrained by model. Product's goal is to serve users and scale up, so it shouldn't dance with shackles on, even if they're golden shackles. For us, whichever model lets the product run faster should be called; we never required the product to use our own model from the start.
But from another angle, the model team does indeed need to support the product in many scenarios. For example, when we make Music Video, the first step requires more accurate music analysis, identifying rhythm, beat points, and other fine-grained information — this is where the model team can come support and make music analysis more accurate. Another example: in video generation, some scenarios work better with our own model, or have lower costs — these can also directly support the product.
So this isn't simply strong coupling or weak coupling. More precisely, the product first runs at its own pace, and the model provides support at critical junctures: on one hand improving effects, on the other lowering API call costs and improving margins, helping the product scale bigger.
GeekPark: Dual-wheel drive is definitely good, but it's also definitely hard.
Yue Cao: My feeling is that startups find it easier to get dual-wheel drive right. The reason isn't the small team size itself, but that startups more easily have a group of people truly in founder mode. Whether doing business, product, or model, as long as the goals in their hearts align with the company's goals, this thing is easy to push forward.
Conversely, if a person doing models is thinking "I want to make a special model, whether the company does well or not doesn't matter much to me," then their goal is actually only aligned along the model line — this isn't dual-wheel drive, but single-wheel drive.
So what truly determines whether dual-wheel drive can work isn't the formal coexistence of model and product in the company, but whether leaders on both sides believe that having both model and product is more beneficial for the company as a whole.
GeekPark: Specifically, how do you handle the problem of "first use the best model to get the product running, then pull key capabilities back in"?
Yue Cao: In the stage from product 0 to 1 finding PMF, if you bind too tightly with your own model from the start, the validation cycle gets lengthened, which isn't conducive to rapid validation and quickly finding PMF. So our approach over this period has been to first build the product with the best-performing models.
At this stage we don't prioritize cost first, but first look at what state it can run to, whether this output can be delivered, whether it can form a commercial closed loop. After this chain is running smoothly, we then look at what places are worth optimizing and worth pulling back in.
So this isn't about requiring the product to use its own model from the start, but first letting the product run at its own pace; the model team provides support at critical junctures. On one hand making effects better, on the other lowering API call costs, improving margins, helping the product scale bigger.
Trust Is the Deepest Moat
GeekPark: What level has your commercial revenue reached now?
Zake Zhang Zihe: VidMuse launched around mid-January, and in roughly two months reached $10 million ARR, and it's still growing. Basically $200,000+ in weekly revenue, and already quite stable.
On pricing, we currently use subscription plus top-up packs. Registered users get 1,000 free credits to start a project.
GeekPark: What does 1,000 credits mean?
Zake Zhang Zihe: Roughly enough to push a 30-second video project to a relatively later stage.
GeekPark: What about paid conversion rate and average revenue per user?
Zake Zhang Zihe: Registration to paid conversion is roughly 5%-7%. Average revenue per user has consistently been quite high, because users need to subscribe first, then buy top-up packs, and some will eventually directly upgrade to higher-tier versions.
GeekPark: Going further? What capabilities will VidMuse 3.0, 4.0 need to fill in? How will product boundaries change?
Yue Cao: 3.0 or 4.0 should be a more thoroughly unleashed state: users mention a feature that wasn't originally in the product, and it can figure out how to mobilize the resources it has to solve this problem.
This will increasingly rely on more general agent capabilities, especially coding agent capabilities. Because the community will constantly produce all kinds of bizarre needs. You need a capability that can flow along with user needs — users give you a link, a post, a tutorial, you can understand the methods inside it, then implement it. The product will rely less on preset features, and more on flowing along with user needs.
GeekPark: Sounds like future products will increasingly "do by not doing." From a long-term perspective, what is Sand.ai's moat? How do you retain users and precipitate long-term value? I believe it's not just model capability, right?
Yue Cao: One of the biggest problems with AI agent products right now is very poor stability, making it hard to establish trustworthy relationships with users.
So our thinking is: first solve various hallucinations, especially the problem of small hallucinations being continuously amplified in multi-turn dialogue, so users dare to trust you. We hope that when users finish creating, what remains is emotion like "thank you," "good night," rather than being angered and depleted. The first step is establishing a sense of trust.
The second step is making users willing to stay here. A good product should continuously get to know this person, understand this person, understand what they like during use. For example, if a user has clearly said they like Nolan, don't push other directors' styles to them; if a user has said they don't like purple, subsequent scenes, storyboards, and script designs shouldn't go in this direction.
So memory (long-term memory) and trust relationships are the soul of our Video Agent.
This article is reprinted from GeekPark
Edited by Siqi Cao; Interviewed by Peng Zhang, Siqi Cao

