MaHui | Sand.ai Raises Over $100M Across Two Rounds, Existing Investor Source Code Capital Continues to Increase Stake
Source Code Capital led the seed round of Sand.ai in 2024 and was one of the company's earliest institutional investors.


Recently, MaHui portfolio company Sand.ai completed two consecutive funding rounds totaling over $100 million. The investor lineup is notably strong: Look Capital, Lollapalooza Capital, Jiukun Venture Capital, Matrix Partners China, MSA Capital, Sinovation Ventures, Xiang He Capital, Source Code Capital, CASSTAR, Hongtai APlus, Capital Today, Hua Capital, Yunhui Capital, IDG, and Baidu Venture, among other top-tier firms, participated in the joint investment. Xinghan Capital served as the financial advisor for this round.
Source Code Capital led Sand.ai's seed round in 2024, making it one of the company's earliest institutional investors, and has continued to increase its stake in subsequent rounds. In this latest financing, Source Code Capital once again added to its position as an existing shareholder, extending its long-term support for the company.

The following content is republished from Z Potentials:
Behind the simultaneous bets by multiple top-tier institutions on Sand.ai, is this a valuation of current achievements or an early vote on future technological paradigms? The answer likely lies in this startup's technical choices and its key judgments about the future.
1
Betting on Autoregressive, Conquering MoE: A Company That Dares to Define Rules Amid Uncertainty
In the era of large models, a company's choice of technical path reflects the acuity of its future judgment. In the video generation field, Sand.ai is one of the few companies willing to place bold bets and explore fundamentals even when the direction remains unclear. This has given the company a significant lead in technical路线选择 — at a time when Diffusion was still the consensus in video generation, they were among the first to shift research focus to autoregressive architecture, becoming one of the earliest definers of this direction. At that time, Sand.ai founder Yue Cao's judgment was that video is not a pixel generation problem, but a compression problem of spatiotemporal and physical laws. Compared to Diffusion, autoregressive approaches hold greater potential for real-time interaction, long-term prediction, and world understanding.
This judgment has been validated by tangible results. Their Magi-1 autoregressive video world model, released in early 2025, achieved absolute leadership on Physics-IQ, the physical realism benchmark proposed by Google-DeepMind, even surpassing Nvidia's latest flagship world model cosmos3-super, and far exceeding other pure diffusion models like Sora2.

Moving from content generation to world understanding, autoregressive has become an important starting point for Sand.ai's bet on the future.
But the real world never exists as a single modality. Human cognition of the environment essentially comes from the synchronous fusion of visual, auditory, motion, spatial, and other information. Therefore, relying solely on video pixel training inevitably yields limited information.
Under this thinking, Sand.ai launched its native audio-video co-generation model in late September 2025, incorporating sound signals into a unified modeling framework — among the earliest teams in China to deliver audio-video co-generation. They found that when the model jointly models both, sound can help generate more realistic visual details, and visuals can likewise assist sound generation.
The judgment Sand.ai offers behind this is that only through high-dimensional, multimodal joint modeling can the world laws compressed by the model approach the true expression of the physical world, thereby avoiding the "cognitive disconnect" brought by traditional single-modality generation.
By 2026, as model scale continues to expand, new challenges emerge. Video world models simultaneously face a triple constraint of quality, speed, and cost. Breaking through this long-standing "impossible triangle" has become the focal point of next-stage competition.
Sand.ai's latest technical blog elaborates on their recent progress in this direction — to address scaling bottlenecks, they have shifted from traditional Dense to MoE (Mixture of Experts), with multiple innovations in architectural engineering.
First, compared to dense models where all network parameters participate in computation, MoE architecture can dynamically activate only a subset of expert networks based on generation content, continuously expanding model size while significantly reducing training and inference costs. Beyond this, the team introduced a novel routing mechanism to optimize communication efficiency, improve expert granularity, and enhance training stability — specifically addressing challenges of applying MoE to video models. This series of innovations enables the model to achieve better balance among quality, speed, and cost.
Another notable choice is that Sand.ai adopted a single-stream unified architecture rather than the multi-stream architecture common in the industry, proposing to map different modalities including text, image, video, and sound into a unified token sequence, processed by a single Transformer.
Under the single-stream architecture and MoE dynamic routing mechanism, different expert networks automatically learn parameter division and modality collaboration relationships based on input content. This means the model no longer relies on manually preset fusion rules, but can autonomously discover associative structures between different modalities during training.
To maximize efficiency and tackle computational challenges from multimodality and long sequences, Sand.ai has continued investing in underlying infrastructure R&D, with systematic optimizations for long sequences and heterogeneous attention scenarios. For example, the team's Magi Attention and other innovative operators significantly improve training and inference efficiency while maintaining modeling capability, reducing long-context computation overhead. These underlying capabilities are critical to how large a model can scale and how complex the tasks it can handle.
Surveying Sand.ai's technical evolution — from the early decisive bet on autoregressive approaches, to firm commitment to multimodal joint modeling, to current innovations in underlying architecture and operators — its technical choices have consistently anchored on the same ultimate pivot: driving the model to transcend the surface of "content generation" and truly crystallize into understanding and simulating the operating laws of the real world. This capacity for deep compression and reconstruction of real-world states is precisely the key variable affecting the ultimate global competition in world models.
2
Video Generation Is Not the Destination of World Models, But the Most Important "Gas Station" Along the Way
In Yue Cao's view, video generation was never the endpoint.
In recent years, discussion of world models has continuously heated up, yet the definition remains unsettled. Many summarize it as "Predict Next State." Cao agrees that prediction is the core of world models, but remains wary of "humans trying to define what the hidden state is." History has repeatedly proven that every attempt to deconstruct the world with human priors fundamentally underestimates its complexity.
This lesson, the history of large language models has already fully demonstrated. On the road to LLMs, countless efforts attempted explicit modeling of word representations, sentence representations, paragraph and even entire document structures — they were elegant, beautiful, aligned with human intuition about "understanding language," and stage-wise were indeed proven "efficient." But on the truly scalable path, they were all, without exception, killed by the most naive approach: predict next token. In the end, no one could define for the model "what the state of language is."
So Cao offers his true judgment: what should really be predicted is not any human-defined state, but the one thing the world gives you for free and that comes with built-in supervision signal — observation itself. This leads to his further conclusion: directly modeling raw data to build world models may not be the locally most efficient approach, but is likely the most scalable one.
And among all raw observations, what comes closest to the real world? Cao's answer is video.
The evolution of video models is precisely the process of continuously approaching the real world: earliest capable only of generating single images, then learning temporal continuity; audio-visual synchronization gave it the sound dimension; multi-shot generation introduced spatial relationships; future prediction established causal associations; real-time interaction brought closed-loop feedback. Each capability improvement was not artificially stuffing in another "state variable," but letting the model grow its own understanding of space, time, sound, and causality from more complete observations. When these dimensions are unified in modeling, the video model will eventually evolve into a true world model in the full sense.
Of course, viewed more rationally, the endgame of world models is vast, and no one can reach it in one step. At the present stage, video models have already demonstrated commercial viability in short video, short drama production, and content creation markets, allowing the technical momentum of "daily incremental progress" to be monetized. Commercialization and technical evolution are not separate either — real demand brings cash flow on one hand, and continuously generates new user feedback and data to fuel model iteration on the other.
Just as next-token prediction was the path that ultimately won for reasoning, Cao believes that next-frame prediction for embodiment will be the same path: refuse to erect another layer of artificial states above observation, let the model optimize itself.
And video generation is not the destination of world models, just the most important "gas station" along the way to that endgame.
3
What's Truly Scarce Is Not the Model, But the Team That Can Define Technology
As world models gradually move from technical concept to engineering reality, a more fundamental question begins to surface: what kind of team can truly pull this off?
Compared to language models, the complexity of video models and world models is elevated across the entire system. Whether in architectural complexity, data supply, or compute consumption, this is destined to be a path for the few.
Globally, teams truly possessing first-tier video foundation model capabilities number no more than five. Competition has accordingly shifted.
Early on it was competition in single-point capabilities — generation quality, resolution, or speed; entering the world model stage, competition shifts to system capabilities encompassing data systems, model architecture, training efficiency, and product closed-loop. At this stage, what determines victory is the organization itself.
Sand.ai founder Yue Cao was previously co-founder of Lightyears Away, head of research centers at BAAI, and lead researcher at Microsoft Research Asia. His履历 spans both basic research and engineering落地, representing globally top-tier technical strength. One of his representative works, Swin Transformer, has become an important foundational component of vision Transformer architecture, and won the Best Paper Award (Marr Prize, one of the highest honors in computer vision) at ICCV 2021.

In terms of academic influence, his papers have been cited nearly 90,000 times, representing a typical basic research-driven technical background. This background determines that the team leans more toward "problem essence" in architectural choices, rather than short-term engineering optimization.
Algorithm lead Zheng Zhang has equally impressive credentials. Former researcher at Microsoft Research Asia (MSRA), gold medalist in the ACM Asian Regional Contest, and core author of Swin Transformer, winning the Best Paper Award (Marr Prize) at ICCV 2021 together with Yue Cao. Google Scholar total citations exceed 60,000, representing a typical academically-driven technical background.
On the product side, operations and growth lead Jia Wang was one of the seven founding team members of Douyin, having gone through the complete 0-to-1 journey as operations director, and also served as consumer-side operations lead at Minimax; additionally, VidMuse product lead Zake Zhang led CapCut PC端's 0-to-1 product strategy and experience design, also worked on OnePlus camera imaging experience optimization, and has long been active as a video content creator in the Bilibili ecosystem, possessing genuine creator perspective and product understanding.
This combination gives the team two capabilities simultaneously: one end understands how models "learn the world," the other understands how content "gets used." This structure itself is a scarce resource. It requires a team that can define both technical boundaries and product forms.
And this scarcity of capability endowment is likewise reflected in Sand.ai's shareholder structure.
If you observe Sand.ai's shareholder structure, a typical characteristic emerges: it doesn't come from a single type of capital, but rather an overlapping combination of multiple types of long-termist capital. There is industrial capital that has weathered many years in the tech industry, dollar funds focused on frontier technology, institutions that have long invested in hard tech and scientists, investors with deep tech company backgrounds, and numerous individual investors who have built companies through multiple ventures.
The logic behind these funds isn't entirely the same. Some value long-term technical potential and care less about short-term returns; some understand the development rhythm of underlying algorithms better and are willing to wait; some are simply betting on the team itself. So this shareholder base brings Sand.ai not just money, but a cognitive network covering different perspectives and experiences. For a company building long-term technology, this combination is more valuable than pure capital.
When a cap table simultaneously gathers these different types of capital, what is being bet on is often no longer a single product, but a technological paradigm that may change industry structure. This opportunity clearly belongs to only a handful of people, and represents enormous beta.



