A Conversation with HiDream.ai's Tao Mei: Big Tech Is a Machine Gun, Startups Only Have One Magazine

**Half a Step Ahead | Unconverged Problems | Compounding Systems | One Magazine**

Half a Step Ahead | Unconverged Problems | Compounding Systems | One Magazine

Produced by | AI Nao

01.

In 2026, the AI industry is cycling back to a familiar pattern.

Following the IPOs of MiniMax and Zhipu AI, investors are once again chasing model capability. The pendulum that had swung toward commercialization and revenue over the past two years is reversing. "Model is productivity, model is product" — Mei Tao has noticed the industry returning to 2023.

Earlier this month, HiDream.ai's new commercial image generation model HiDream-O1-Image-1.5 ranked in the global top three on Artificial Analysis, the authoritative independent AI benchmarking platform. HiDream became the world's second-highest rated image foundation model company — and China's highest — trailing only OpenAI. Just in May, its open-source model HiDream-O1-Image had topped the same platform's open-source rankings worldwide.

  • June 9: HiDream-O1-Image-1.5 ranks in Artificial Analysis top 3

On the capital side, the company announced in April that it had raised over 500 million RMB, and within two weeks closed another round with participation from Shenzhen Capital Group, GP Capital, Caixin Capital, and Fuju Capital.

Meanwhile, as video generation, multimodality, and embodied intelligence continue to heat up, "world models" have been pushed to the center of the multimodal foundation model narrative. But unlike many peers, Mei Tao has remained restrained toward this concept. He prefers to see it as a direction that is gradually being approached, not yet arrived.

"We don't believe the industry has fully achieved world models today," he says.

02.

Looking at HiDream.ai's trajectory over the past three years, this restraint is hardly surprising.

Twelve years at Microsoft Research Asia, including the experience of releasing TGANs-C, the world's first video generation model, convinced Mei Tao that enormous opportunities remain in model innovation. Five years leading AI business and industrial deployment at JD.com taught him something else: technical leadership does not equal commercial success, and model capability does not automatically translate into market advantage.

This is why HiDream.ai, from day one, never positioned itself as a pure model company.

In Mei Tao's view, model capability itself does not automatically form a moat. Startups need sustained R&D investment, but must also enter real industries as quickly as possible — gaining data in scenarios, validating capabilities, and establishing commercial loops. HiDream.ai's positioning is to do both models and applications, with both driving each other.

This also determined that HiDream.ai was never a company focused solely on base models. Mei Tao summarizes its business model as "1+1+3": at the bottom, self-developed multimodal foundation models; in the middle, the Token Hub platform, which standardizes model capabilities into callable outputs; at the top, three specific scenarios — FrameZan for AI film and television, HiBurst for commercial marketing, and vivago for content creators.

These choices appear to span wide areas, but they follow the same underlying logic.

"Big tech is a machine gun with endless ammunition. A startup has one magazine — every round fired is one less."

In Mei Tao's view, the most dangerous thing for a startup is not falling behind on a single model benchmark, but being dragged into a war defined by big tech. Compute, capital, and talent reserves put the two sides in entirely different weight classes. A startup's resources are extremely limited; every choice means abandoning nine others, and every investment must point toward greater future possibility.

Therefore, rather than defining HiDream.ai as a video model company, an Agent company, or a pure B2B company, it is better understood as a platform company expanding application scenarios around a native omni-modal model foundation. Models, scenarios, commercialization, and embodied intelligence are never separate business lines in Mei Tao's eyes, but interlocking gears serving one goal: ensuring limited resources compound rather than deplete.

「 Conversation with Mei Tao 」

A Startup Has Only One Magazine

It Must Be Fired Toward the Future

AI Nao HiDream has made many choices that look non-standard over the years. Including now — you're not placing yourselves within the mainstream world model narrative. How do you judge what a company should and shouldn't do?**

Mei Tao I've always believed the most important capability for a startup is not addition but subtraction.

Opportunities in AI today are abundant; new hot directions emerge almost monthly. But startup resources are finite — you can't do everything. Big tech is a machine gun with endless ammunition; a startup has one magazine, and every round fired is one less.

So our criteria are actually simple: can this thing generate sustained accumulation?

If it's just chasing trends, capturing one-off opportunities, even if effective short-term, it's hard to form lasting advantage. We care more about whether a direction can simultaneously bring customer, data, and capability accumulation.

I've always believed in this: root first, then grow. Establish your position in one domain, then expand boundaries. Otherwise you easily fall into the competitive rhythm defined by big tech.

AI Nao But entering competition in the same arena as big tech is often an easier way to be seen by capital and markets.**

Mei Tao Many people feel that entering the main battlefield proves your strength. But the problem is, big tech most wants you to enter the main battlefield. Because that's where it excels most.

If the final comparison is compute investment, capital scale, and organizational capability, then the outcome was essentially decided from the start. Leading video generation models like Sora and Seedance are backed by long-term, ultra-large-scale compute, data, and engineering investment.

Startups cannot infinitely compete with big tech in a war of attrition. Our ammunition is extremely limited, so we must choose more efficient paths.

  • Mei Tao, founder of HiDream.ai (HiDream)

Must Run Half a Step Ahead Before Consensus Forms

AI Nao Where do you choose to fire your limited ammunition?**

Mei Tao I care about whether, with limited resources, I can find opportunities that haven't fully converged yet.

For example, multimodal models are different from large language models. LLMs today have basically converged on the Transformer architecture; it's mostly a competition over data and compute. But multimodal is far from that stage — model architecture, training methods, data organization all still have many possibilities.

In moments like this, startups have opportunity.

Because you may not have more resources, but you can be faster than big companies.

We started working on UiT (Unified omni-modal Transformer architecture) at the end of last year; this is essentially the same logic. Since we can't compete on resources, we strive to move one step ahead before a technical direction forms consensus. By the time everyone sees the direction, the window may have already closed.

So I've always felt that what startups should really compete for is not the same position as big tech, but half a step earlier. Making innovations at the underlying architecture first, proving out new technical paths first.

AI Nao What's the specific difference between UiT and what we generally call multimodal?**

Mei Tao Simply put, most multimodal models today still essentially treat images, video, and text as different things. Each has its own model, trained separately first, then somehow aligned to translate between each other.

But UiT is different. Traditional multimodal mostly trains text, images, and video separately, then connects them through alignment modules. UiT, from the architectural design stage, puts different modalities into the same system to learn together. We have an internal phrase for it: "childhood sweethearts" — these modalities don't meet each other after growing up, but grow up together in the same model from the start. This means the model learns not just one kind of generation capability, but deeper structural relationships between modalities. Only on this basis is it possible to truly move toward Any to Any: arbitrary input supporting arbitrary output, and further possessing the capabilities a world model requires — understanding, generating, and predicting different states of the real world.

AI Nao Why do you feel this is worth investing in?**

Mei Tao **Because it corresponds to a problem that hasn't converged yet.

Much past model innovation has essentially been optimization within existing paradigms — more parameters, higher training efficiency, lower inference cost. What attracts us about UiT is that it attempts to answer a more fundamental question: how should a model understand the world?

Today's large models can already generate very good images and video, but often it's more like statistical fitting. It knows what things frequently co-occur, without necessarily truly understanding the relationships between them.

If a world model truly exists in the future, it won't just generate the world — it will understand the world, generate the world, and predict the world.

I don't believe the industry has reached that point today, but I think UiT represents a path we believe could lead to world models — using a native omni-modal architecture to put visual generation and world representation into the same framework.

For a startup, this is the kind of opportunity I was talking about. It hasn't formed industry consensus, hasn't been proven correct, but if it holds, it could open new space.

AI Nao You mentioned "understand the world, generate the world, predict the world." But many companies have already defined themselves as world model companies — why do you think everyone is still far from it?**

Mei Tao **I think the biggest problem here is that use of the term "world model" is more a concept than an established capability stage.

Many models today can indeed do very strong generation — longer video, more stability, more natural multimodal alignment. But this is more about getting stronger at "generating content," not a qualitative leap in "understanding the world."

To put it more intuitively, current models are more like learning "what the world looks like," not "why the world is this way." The difference between these two is actually huge.

The former can be fitted through massive data; it's still essentially statistical fitting of surface appearances. The latter requires the model to establish stable internal representations of physical laws, spatial relationships, causal chains, and dynamic changes in the real world, and to make reasonable inferences and predictions under new conditions. A true world model cannot just "generate a world that looks reasonable" — it must understand how the world operates, why it operates this way, and be able to reconstruct it under new constraints.

So I believe there is still distance from a true world model in this sense. Not that there's no progress at all, but that the key threshold hasn't been crossed.

The industry today is more like approaching that direction, but hasn't entered what could be called a converged stage.

From Content Generation to World-Building:

How Model Capabilities Grow Into Closed Loops

AI Nao You've consistently emphasized entering real scenarios, getting feedback, building closed loops. But many AI companies, once they go deep into industries, easily become project-based and delivery-based like the previous generation of SaaS. How do you judge whether you're accumulating capabilities or being led around by customers?**

Mei Tao Our criteria are actually very simple. If a need only solves one company's problem, it has no value to the model. Only common problems representing a category of scenarios do we consider doing.

Many companies, once they enter an industry, get pulled around by customer demands. Whatever the customer wants, they do — ending up with massive amounts of customized projects. Short-term revenue may look good, but capabilities don't accumulate. From day one of founding, we've had one principle: we don't do this.

So we don't do heavy customization, and we don't do heavy integration. We'd rather be the party that others integrate, than the one doing the integrating.

AI Nao So when you choose cross-border e-commerce, content e-commerce — you're not essentially looking at the industries themselves?**

Mei Tao What we value isn't the industry, but the feedback mechanism.

Many people ask why we do cross-border e-commerce and not gaming, not other industries. The core reason is actually simple: this industry can continuously generate feedback.

For example, content e-commerce clients produce content every day, run campaigns every day, see conversion results every day. This means the model gets new data and new validation every day.

We now serve over 40,000 clients, but HiDream's team is only 200-plus people. With a traditional enterprise service model, this would be nearly impossible. But because we serve massive standardized needs rather than massive customized needs, the entire system can operate at scale.

An important standard for platform companies: if you pull out a batch of operations staff to do something else, and company performance doesn't immediately drop, that shows the system is starting to have platform attributes. If growth must depend on more and more people, it's still essentially a delivery logic.

So when we choose a scenario, the most important thing isn't how big the industry is today, but whether it can continuously bring us customer, data, and model capability growth.

AI Nao You don't see models, commercialization, and scenarios as separate things.

Mei Tao Many people oppose technological innovation and commercialization, as if people doing tech shouldn't think about customers, and people doing commercialization don't need to understand models. But startups don't have this luxury. Because every investment a startup makes must form the foundation for the next investment.

Model capability brings customers, customers bring data, data in turn improves models — this should be a continuously strengthening cycle.

If this cycle holds, the company gets stronger and stronger. If it doesn't hold, then no matter how advanced the model or how good the revenue looks, it's essentially just consuming resources.

So I never see models, commercialization, and scenarios as three business lines. To me, they're more like a set of interlocking gears. Models determine what scenarios you can enter, scenarios determine what data you get, and data determines how far models can go.

The most important thing for a startup isn't expansion, but accumulation. Because big tech wins through scale; startups can only win through compounding.

AI Nao For the past 3 years, HiDream has been doing visual models and content generation. But recently you've started doing embodied intelligence-related data business. Many people would see these as two completely different directions — how do you see the relationship between them?**

Mei Tao If a true world model is ever realized, it ultimately cannot stay only in the digital world. Today's video generation is essentially learning the visual laws of the world, while robots face a more complex problem. They need to know not just what the world looks like, but how it moves, how it changes, and what results from taking action.

So embodied intelligence is actually forcing models to further understand the real world. The model doesn't just know what the world "looks like," but how it "changes," and what results might follow from an action. We're doing synthetic data now because real robot data is too expensive, too scarce, with collection cycles too long and scenario coverage too limited. Many capabilities can't be slowly accumulated from the real world. But if we can supplement robot training with more scenarios, actions, and feedback samples through high-quality simulated generated data, there's opportunity to accelerate this process.

From this perspective, multimodal understanding and embodied intelligence, I believe, are actually solving the same problem: extending multimodal generation capability from "generating content" to "building world" — enabling models to simulate and shape the real world on the basis of understanding physical laws and predicting state changes.

AI Nao How do you hope people define HiDream in ten years?**

Mei Tao What I pursue isn't people remembering which hit model we made, because models will always be replaced. I hope that in ten years, HiDream.ai will be a globalized, platform-based company centered on a native omni-modal world model foundation, consistently doing enterprise services well.

What I hope more is that when the industry looks back on this wave of AI development, it will find that we made serious, solid explorations in some directions where there was not yet consensus.

Image sources | Generated by HiDream-O1-Image-1.5, provided by interviewee

Join the Community