MaKe | Dark Current Dialogue with MaHui Member Yue Cao of Sand.ai: Farther from Sora, Closer to the Endgame


- This article is republished from An Yong Waves (暗涌Waves), by author Yu Lili.
In the history of China's large model development, "Lightyears Away" is undoubtedly the most dramatic episode.
It was the earliest "China's OpenAI" in the public consciousness. It began in February 2023 with Huiwen Wang's rallying cry for talent. Then, due to Wang's illness, it ended abruptly. The whole affair lasted barely four months.
The two co-founders Wang recruited — "one with an infrastructure background, the other with an algorithms background" — were Jinhui Yuan, who later founded SiliconFlow, and Yue Cao, who founded Sand.ai.
Such a tumultuous opening caught many off guard. Cao likened himself at the time to "a tree in winter" — "no visible changes on the surface, but roots growing furiously underground." A few months later, in January 2024, he established Sand.ai.
That year was one of clamorous activity in video generation: Sora detonated the field early on, and Keling AI made a stunning debut mid-year. And whether ByteDance, creating and understanding 3D structures. From this, we see new possibilities for AI for Science.
As for Sand.ai, apart from being cited by venture capital "queen" Kathy Xu last July as evidence of her still-active investments, it has remained completely silent.
From this perspective, Sand.ai is undoubtedly the latecomer.
But in Cao's view, the lateness stems from their choice of a harder, more fundamental path — one fundamentally different from Sora's.
On April 21, Sand.ai officially launched Magi-1. Unlike Sora, which chose the DiT (Diffusion Transformer) route, Sand.ai integrated diffusion models with the higher-ceiling autoregressive approach.
Grounded in first-principles thinking about video data, Cao believes that in video generation — where the technical path has yet to converge — autoregressive (temporal autoregression) may be the solution closer to the endgame. He stated: "Magi-1 is our first attempt to bring autoregressive video generation back into the mainstream. This will be an interesting beginning."
In this direction, he believes Sand.ai is not late at all, but rather arrived first.
Earlier, Cao had worked in the star-studded Visual Computing Group at MSRA (Microsoft Research Asia), which produced Shaoqing Ren, Kaiming He, Xudong Cao, Xiangyu Zhang, Jifeng Dai, and others.
In 2021, as one of four core first authors, he published Swin Transformer, which won that year's ICCV Best Paper Award (the Marr Prize). Afterward, driven by thinking about organizational forms closer to a "China's OpenAI," he joined BAAI (Beijing Academy of Artificial Intelligence).
Then came the tsunami of ChatGPT. This post-90s researcher, who before age 30 had never considered entrepreneurship and idolized Ilya Sutskever, was swept into the real commercial jungle.
Recently, we sat down with this young founder to talk about his adventure.
A video generated by Magi-1.
The conversation follows:
A Different Story from Sora
Waves: After Lightyears Away was acquired by Meituan in late June 2023, you largely disappeared from public view. What were you doing then?
Yue Cao: I left around August or September. I felt I needed time to think through what to do next. A friend described me then as very much like a tree in winter — no visible changes on the surface, but roots growing furiously underground. Overall, I really enjoyed that process.
Waves: Why did you choose to start your own company after leaving Lightyears Away?
Yue Cao: Pursuing extreme personal growth is a very fundamental drive for me. Entrepreneurship felt like that kind of option. For a long time, I didn't know exactly what I wanted to do, but I was very clear about what I didn't want.
For example, I didn't want to stay in a stable, predictable system where I could see the end from the beginning. Because on that track, there wasn't the kind of growth I wanted.
Even though entrepreneurship is incredibly tough, founders generally gain far more refinement and growth than ordinary people. My time at Lightyears Away also deepened my understanding of entrepreneurship. I realized this was the career and state I'd been pursuing all along.
Waves: In December 2024, Sora released its demo. Did that influence your decision to enter video generation?
Yue Cao: Not directly. I decided on the video direction around November 2023. Before that, I'd looked at many directions, including companionship, agents, coding, and so on.
I ultimately chose video generation because it's a direction with both very high technical ceiling and very high commercial ceiling — "a long slope with thick snow." If you think from first principles, AGI also cannot do without compression of video data.
Waves: But in the video generation track, compared to earlier entrants with product launches already, this timing makes you a latecomer.
Yue Cao: It's actually hard to put everyone in the same race. I think in the direction of using autoregressive technology to compress video data, we are the pioneers. This path is not equivalent to video generation; video generation is just its first clear application scenario.
Waves: At the time, more companies were probably thinking about replicating Sora, while you were telling a completely different story.
Yue Cao: First, Sora's current form may not even be OpenAI's goal in this direction — it could be a smokescreen.
Thinking from first principles of technology, Sora's technical route itself has clear problems, and its ceiling is likely not very high — not sufficiently scalable.
When the generation quality ceiling of the Sora approach isn't high enough, its contribution to AGI itself is limited. When the company was first founded, we had intensive discussions to find a solution closer to the endgame.
Waves: Did you find an answer?
Yue Cao: We believe video generation needs AR (autoregression). From our perspective, this is a solution closer to the endgame.
Magi-1, which we just released, is a beginning — the first milestone of the AR (autoregression) route. It proves that the AR approach is feasible, and its effects reach first-tier levels among video generation models on the market.
For the entire community, the AR route now has a model that completely doesn't lose to pure diffusion routes. I find that quite interesting.
Waves: Why are you convinced that the autoregressive route is necessarily the solution closer to the endgame?
Yue Cao: This is our intuition about the technical direction. We believe video is ultimately causal in the temporal dimension. Just like language models, you can only read text from top-left to bottom-right; no one reads backwards. Video is the same. Many physical laws are essentially functions that change over time.
But Sora doesn't have these settings. In early Sora or Sora-like solutions, when a person walks, you often get cases like left leg, left leg, right leg, right leg — instead of: when the left leg steps in the previous second, the right leg should step in the next. This is because the model only learned temporal correlation during training, not temporal causality.
Temporal causality is one dimension; another is that I believe the autoregressive route is more scalable.
Waves: What differences in experience does this temporal causality directly bring?
Yue Cao: From the most fundamental "model capability" perspective, it should be open-ended and a posteriori.
From the product perspective, it unlocks some natural product characteristics. For example, breaking free from duration limitations, and enabling more granular control over the time dimension — Sora's storyboard attempts this, but with poor results because of inherent model limitations.
Waves: Though at the time, the DiT architecture having a lower ceiling while the AR architecture having potential to break through that ceiling was somewhat industry consensus.
Yue Cao: At the time, the importance of this direction was relatively consensus, but there wasn't a public, widely recognized solution. And when to do it, how to do it, and how far to explore — these make huge differences.
Some companies might treat making Sora as step 1 and AR as step 2, but in terms of directly incorporating the autoregressive route into video generation tasks, we were very early. In fact, we've already traversed much that other companies haven't had time to do.
Waves: From afar, as early as December 2023, Google Research launched a video generation model based on autoregressive architecture. More recently, OpenAI's GPT-4o, DeepSeek's released multimodal large model, and ByteDance's published papers all mention autoregressive architectures or models. How are these different from what you've built?
Yue Cao: All of these involve autoregression, but at very different granularities. With 4o, in my view, the more critical thing is making image generation and language models into one model, unifying so-called multimodal generation and understanding.
In autoregressive architecture, video generation requires causality in the time dimension. This concept may have occurred to many people or been mentioned, but actually putting it into practice, at relatively large scale, making it work, and with good results — we haven't seen that yet.
Waves: So this wasn't an easy decision to make?
Yue Cao: People and organizations who can make judgments about technical routes are still relatively few — OpenAI already handed you the homework, yet you choose another path.
Often, following is easy because you can more easily convince others and the whole organization. Not following is much harder; everyone needs to realign, build shared understanding, and because it has more challenges, it will also be much slower.
Identifying the True First Curve
Waves: What do you think of Keling AI 2.0's release? It's said to have systematically studied scaling law characteristics of video generation DiT architecture. And when Keling AI first launched last June, did it shock you?
Yue Cao: Studying and understanding scaling laws is the first step; the key is still which path — DiT or AR — is more scalable. We're still in early days, with a long road ahead.
At the time, Keling AI came out so quickly with quite good results — that was surprising, since other companies basically achieved that around September or later.
Waves: What do you think might be the secret to its speed and efficiency?
Yue Cao: There are many factors; it's hard to directly attribute.
Generally speaking, they had accumulated capabilities in both engineering and organization. Video data processing is quite different from language models — its storage overhead, processing overhead are all much larger. For a startup, engineering support is really needed. Organizationally, at that point in time, they were also very united, like a startup itself.
Waves: In your view, what are the differences in approach between Kuaishou's Keling AI and ByteDance's Dreamina?
Yue Cao: From the product side, Keling AI may be more focused on making a good tool — model as product — while Dreamina wants to build a new type of content platform, or an AI video version of Douyin.
Waves: Currently, looking at overall funding market activity, this track's volume has decreased quite a bit. Not only have major platforms secured more advantageous positions, but some large model startups are also pushing in this direction. Do you think startups still have a chance to break through?
Yue Cao: Long-term, whether in technical ceiling or commercial ceiling, I think this direction can absolutely support a good startup. Midjourney previously achieved $200 million to $500 million annual revenue with high gross margins. I think the entire market this year can completely reach that level, or even higher.
If there's a depressed state, it's because model quality isn't good enough, not in the first tier. From another perspective, if a startup achieved Keling AI's position, would it lack money?
Also, for startups, I think the main thread is very important. Every decision determines what kind of company you ultimately want to become. Founders must first identify what the true first curve is, then go deep enough in that direction, iterate continuously. One should ask whether so-called second curves, third curves exist because the first curve was never stable.
In this era, if you don't have sufficient judgment about technology itself, don't know which direction to develop toward, it's hard to feel secure.
Waves: How's your sense of security? What do you rely on for strategic choices?
Yue Cao: I'd say I'm strategically optimistic, tactically pessimistic. Strategic choice often depends on the team and CEO's thinking about their own endowments and long-term chances of winning. For example, are you fundamentally a company that can make great models, or one that's more agile at productization and commercialization?
My experience determines that I'm better at model and algorithm-related problems. Our team was also relatively model-side-oriented from the start. These give us a relatively short decision chain on technical matters — this is our strength.
Waves: And the weakness?
Yue Cao: I'm a first-time founder. I haven't done products, haven't done commercialization — this part definitely requires time to learn.
Waves: So you tend toward a strategy of amplifying your strengths.
Yue Cao: Looking at DeepSeek, you could argue it's sometimes strategically conservative. I think the key is that in execution, it has very full and accurate self-awareness of its capabilities — this is most critical.
You should grasp what's most important to you. At the current stage, I believe technology is still most important.
Waves: But not all startups have DeepSeek's real-world conditions. Many startups can't focus precisely because of more immediate commercialization pressures.
Yue Cao: Our product was just released; we'll consider commercialization later. Probably first consider going overseas to some high-value regions. Overall, we'll still do more model-as-product rather than being too productized. Model products can actually be used in many scenarios.
Also, unlike language large models, video generation is relatively close to commercialization. Runway, Keling AI, and Hailuo AI all have very high revenue.
The reason is simple: human productivity in the video category is truly terrible. Though current quality is still not good enough, hallucinations are quite serious, and many complex motions are still unachievable — even in this state, it can already satisfy many scenarios.
For example, if you draw 20 times and get one good result, it's still far more convenient than traditional filming where you need to set up scenes, prepare props, and house and feed actors. And video penetration is an order of magnitude higher than text. If its production cost and cycle can be reduced, the value created for the entire market is very considerable.
Waves: Both Robin Li and Allen Zhu have previously expressed varying degrees of pessimism about the video generation direction.
Yue Cao: I'm not sure how others see it, but I think if you think toward AGI, video may be another critical data type beyond language. It's more easily obtainable, sufficiently rich, sufficiently diverse, and the information itself is relatively self-consistent.
It may be some kind of connection between the virtual world and the real world, while language leans more toward the virtual world.
Moreover, language models have developed to a relatively late stage, while video is still at a relatively early stage. I think the true future video foundation model, or in the long term the so-called "world model," will emerge in this direction.
3
Some Philosophy of Doing Things
Waves: In this fragile era, why did you name the company Sand.ai?
Yue Cao: The main element of sand is silicon. We carbon-based humans are now essentially at the frontier of silicon-based existence, and in speculation, silicon-based life forms all feed on sand.
Waves: In 2023, what was the catalyst for Huiwen Wang pulling you into entrepreneurship? What did you mainly do during those four months?
Yue Cao: Everyone was attracted by the grand vision of AGI. During my months at Lightyears Away, I mainly focused on recruiting.
Waves: What was your direct impression of him?
Yue Cao: This person is really strong. He can articulate very sharp perspectives, and knows in what scenarios they'll work. When he wants to present a point, he can give you a 10-minute, 40-minute, or even 1-hour or 4-hour version based on your situation and comprehension level at that moment.
Often the points he tells you require repeated pondering and chewing, understanding them in different states. It makes you feel that much methodology truly comes from practice, not from books.
Waves: For example?
Yue Cao (after thinking for a long time): Around 2021, I had been studying OpenAI and DeepMind for a long time, and when meeting people, I would often ask why China hadn't produced such organizations.
The first time I met Lao Wang, I asked him too. His perspective was: because China wasn't wealthy enough before.
Waves: Do you agree with this answer? For some highly profitable major companies, this doesn't quite hold.
Yue Cao: Before 2024, compared to Silicon Valley, I think this statement held true.
Of course this may be one side; the other side may be that we got wealthy too fast. The psychological state hasn't adjusted, so there's also the saying "take a generation."
But from the beginning of this year, after DeepSeek and several other companies emerged, there have clearly been very different signals. As Wenfeng Liang said, we just need some facts and a process. We're already in that process.
Waves: Why were you particularly curious about OpenAI and DeepMind at that time?
Yue Cao: I was somewhat confused myself then. Though we had done some work, there was clearly an essential gap compared to more impressive things.
In 2020, the United States produced AlphaFold2, GPT-3, and DALL-E, CLIP which seemed less influential than GPT-3. When I carefully studied OpenAI's previous work, I felt these people's way of working, thinking, and organizational form were quite different from ours.
Domestically, it was still generally paper-driven, while they were some kind of more organized research. They didn't pursue whether methods were novel, but pursued solving essential, important, influential problems. When I discovered this gap, I couldn't stay at MSRA anymore.
Waves: Was this why you decided to join BAAI in 2022?
Yue Cao: Yes. Many Chinese institutions' organizational forms couldn't escape being paper-driven. But if what you pursue is publishing papers, choosing a niche problem may be easier for publishing papers than a widely watched one.
MSRA was already one step further than most organizations, moving from paper-driven to impact-driven, and BAAI was one step further than MSRA.
When I went, BAAI had just undergone some changes. Early on, it supported some university teachers doing organized research inside; later it recruited internal researchers and built a set of values to do more influential, more exploratory things in an organized way.
Because it's a non-profit institution, you could ultimately open-source the results. At that point in time, BAAI was the organizational form with the most chance of approaching OpenAI.
Waves: Has doing influential things always been attractive to you?
Yue Cao: I joined MSRA's Visual Computing Group in 2018. This group was probably the top domestic group doing deep learning. It previously gathered people like Jian Sun, Kaiming He, Xiangyu Zhang, Shaoqing Ren, Xudong Cao, Jifeng Dai, and so on.
Though I wasn't completely contemporaneous with them — some I met later — I later summarized that we were all doing research with somewhat similar methodologies, with some inheritance of philosophy of doing things.
Waves: For example?
Yue Cao: My later summary is: do the most essential, most critical, most widely watched problems. The most important problems are essentially like the 100-meter dash at the Olympics — you need to hone yourself to approach some limit, build deep and fundamental understanding of the problem, and make experiments sufficiently solid, sufficiently fine-grained, to have a chance at real progress on the truly hardest problems. And important problems are like some kind of lever — once you have genuine progress in them, they create enormous impact.
Waves: Before 2023, did you ever imagine being this close to business? Your idol then should have been a scientist, not an entrepreneur.
Yue Cao: It was Ilya. This is easy to understand — he was very close to my direction. In deep learning, he's also one of the few people who can connect the important nodes across the entire AI era, and has produced real, enormous value for this world.
Waves: Is this also the meaning of your entrepreneurship this time?
Yue Cao: In recent years, one of my biggest realizations is that life is quite short. For me, business is a lever; the goal is still to produce some interesting value for this world.




