Hidream.ai Raised 2.1 Billion RMB in Three Months to Push AI From "Generating Content" to "Generating Worlds"

More Than 20 Institutions Scrambling to "Get on Board"?

Over 20 Institutions Scrambling to "Get On Board"?

👦🏻 Author: Ms. Yi

🥷 Editor: Koji

🧑‍🎨 Layout: NCon

On July 20, at the World Artificial Intelligence Conference (WAIC), Hidream.ai unveiled vivago R1, an uncapped-duration multimodal content creation agent. Users simply upload a few photos, add a brief text description, and can generate minute-long videos.

Less than three days later, the company announced it had closed a 1.5 billion yuan Series C round. Over the past three months alone, it has completed three consecutive funding rounds totaling over 2.1 billion yuan, achieving unicorn status.

How does a three-year-old AI company command such concentrated bets from over 20 top-tier institutions at this moment? What's behind the 2.1 billion yuan raised in three months, and how does its latest product, vivago R1, actually perform?

When AI Learns to "Think Like a Director"

AI video generation has a well-known pain point: most tools can only produce 15-to-30-second clips. Extend the length, and characters warp. Want to adjust a character or plot point? You're usually starting from scratch.

vivago R1's first impression is its attempt to solve these problems through "long-horizon reasoning." It functions more like a multimodal agent that can autonomously plan tasks, allocate resources, and drive creative production forward — rather than a black-box tool where prompts go in and videos come out.

Once you start using it, an intuitive impression emerges: it has turned creative methodology into product. Many overseas content creators have also given vivago highly positive feedback.

In vivago R1's design, Skills are no longer mere function buttons but productized encapsulations of creative methods. It compresses what once belonged exclusively to professional creators — script breakdown, shot organization, asset inheritance, style control, and generation sequencing — into a directly callable methodology for users.

Users don't need to understand model parameters or workflows. They only need to know whether they're making a product ad, a narrative short, or branded content. The Skill translates that goal into an "operating protocol" the Agent can execute.

SkillHub is the organizational system for these methods.

It translates abstract AI capabilities into perceptible creative scenarios, letting users move quickly from "what I want to make" to "I can start now."

To date, SkillHub has accumulated hundreds of Skills validated through real projects, spanning everything from cross-border e-commerce viral content generation to film crew storyboarding to individual creator drama-channel operations.

Supporting this Skill system is vivago R1's hierarchical memory and multi-Agent collaboration architecture. Here's the most intuitive way to explain how it "works like a person":

• Top layer: Orchestrator.

It acts as a head director, only retaining final decisions — what are the character settings, what's the narrative throughline, which choices have been confirmed. Process noise gets filtered out here; the system doesn't "hold grudges" over a failed generation.

• Middle layer: Specialist Agents.

Resident here are the Screenwriter Agent, Director Agent, Designer Agent, Editor Agent — each with independent memory, handling their respective domain. The Screenwriter Agent tracks story progression, the Director Agent maintains shot-language consistency, the Designer Agent guards visual-style unity.

• Bottom layer: Execution Agents.

Responsible for segmented parallel generation. Each execution sub-Agent handles only one segment, then exits; if one segment fails, only that segment retries, without affecting overall progress.

At WAIC, Hidream.ai founder and CEO Tao Mei summarized this mechanism's operating logic as: memory lightens layer by layer, final decisions rise layer by layer. This maps tightly to real creators' intuition — a director doesn't obsess over every detail of Scene 1 while shooting Scene 2. What's confirmed floats up as final decision; what's discarded sinks to memory's bottom.

Beyond architecture, vivago R1's interaction mode also feels "natural" to creators. It offers two views simultaneously: Chat and Canvas.

Chat is the task-progression view, handling user intent input, asset assembly, and the "timeline" of sustained collaboration between user and Agent. Canvas is the project-context view, reorganizing key content generated during creation into a browsable, judgeable, reusable results space.

This design lets users "create while chatting" in Chat, while maintaining holistic project oversight in Canvas.

After the video is done, one notably user-friendly detail: it can directly bind TikTok, YouTube, and other overseas accounts for one-click distribution. No more "export file — open platform — re-upload — fill in info — adjust again" rigmarole. For creators, this "smoothness" is often what determines daily output versus abandonment.

From SkillHub's method guidance, to hierarchical-memory multi-Agent collaboration, to Chat-plus-Canvas dual-view workflow, to one-click distribution closing the loop — vivago R1 attempts to build a "creation Agent" that, guided by user intent, completes the full journey from idea to publication. A core trend behind this: the second half of video generation needs not just better models, but Agents that can think.

Over 20 Institutions Scrambling to "Get On Board": Hefei Bets on Another Unicorn

On July 23, Hidream.ai announced its 1.5 billion yuan Series C. Beyond the amount, what's more notable is the "speed" and "investor lineup."

Three rounds in three months, totaling over 2.1 billion yuan. In April, Hidream.ai had just announced a 500 million yuan-plus round, bringing in Oriental Fortune Capital, Anhui Provincial Investment Group's provincial industry investment arm, Fenghua Capital, and others. In May, it secured another round from Shenzhen Capital Group, GP Capital, Caixin Capital, Fuju Capital, and others. By July, the 1.5 billion yuan Series C closed swiftly.

On the investor roster, Hidream.ai has assembled a diverse, strong shareholder structure of "national team + local state capital + industrial capital + top financial institutions." Lead investors include National Social Security Fund Sichuan Zhenxing Sci-Tech Innovation Fund, ICBC Capital, Hongyi Asset Management, and Dunhong Capital; follow-on investors include Xiamen ITG Capital, Shanghai Film New Horizons Fund, Hubei Yangtze River Industry Investment Group, Huace Film & TV, Hangyuan Capital, Chuangyunhai Capital, Huafu Investment, Yuhang Financial Holdings, BOCOM Capital, Ruosong Fund, and others; existing shareholders Hefei Industry Investment, Oriental Fortune Capital, GP Capital, Jinhua Financial Investment, Zhongzhechuang, and Caixin Capital continued to add to their stakes.

Behind this round, the National Social Security Fund — as patient capital — has entered. This signals an important strategic deployment by the NSSF in the native omnimodal world model direction.

On the industrial capital side, beyond Huace Film & TV and Shanghai Film New Horizons Fund entering in the C round, Hubei Yangtze River Film Group had already invested in the A round. The simultaneous presence of three film-industry capital players is unusual in AI large-model company financings. Huace Film & TV has already established an AIGC special fund and an AI video investment fund.

Among all investors, Hefei Industry Investment merits particular attention.

As a company headquartered in Hefei, Hidream.ai's funding history is deeply intertwined with Hefei state capital. As early as 2024, the company received a Series A led by Hefei Industry Investment and Anhui Provincial AI Mother Fund; this year, both institutions continued to follow on, with Anhui Provincial Industry Investment Group, Hefei High-Tech Investment, Xingtai Group, and others also joining the shareholder register. With this round's entry of Chuangyunhai Capital, Hangyuan Capital, and other Hefei-linked forces, Hefei-background institutions now essentially cover Hidream.ai's Series A, B, and C rounds.

This represents an extension of the "Hefei model" into artificial intelligence. Over the past decade, from BOE to Changxin Memory Technologies to NIO, Hefei state capital's playbook hasn't been chasing short-term trends, but rather pre-positioning around key industrial chains, using anchor projects to drive industrial clustering.

Today's bet on Hidream.ai follows the same logic: AI is becoming a new round of industrial infrastructure, and visual generation — along with native omnimodal large models — is among the most application-spillover-rich directions.

This year alone, Hidream.ai has repeatedly appeared as Hefei's benchmark enterprise in artificial intelligence, particularly in the world model domain, at national-level exhibitions including the China Science and Technology Fair and the Chain Expo. From computing infrastructure to AI application ecosystem, Hidream.ai is becoming a key piece of Hefei's AI map.

If Changxin Memory Technologies corresponds to memory chips, and NIO to new energy vehicles, then Hidream.ai's native omnimodal large-model track may correspond to Hefei's next "industrial calling card" — generative AI infrastructure.

From "Generating Content" to "Generating Worlds"

Behind this concentrated fundraising, Hidream.ai wants to reinterpret the relationship between "AI and the physical world."

In his WAIC 2026 keynote "Toward World Models: Native Omnimodality Driving Agent Capability Leap," Tao Mei systematically laid out Hidream's judgment on the next phase of AI evolution. He posed a core proposition:

AI large models are moving from "understanding the world" toward "changing the world." The key to reaching this goal is the dual advance of native omnimodal world models and agents.

In his view, while "world model" is frequently mentioned, industry routes aren't unified: some start from 3D spatial modeling, some approximate world simulation through long video generation, others follow the embodied intelligence route to solve motion planning and physical interaction.

Hidream's choice is the native omnimodal (UiT) route. At the most fundamental level, it encodes "the operating rules of the world" into the model's DNA, letting the model understand vision, language, action, and spatial relationships in a unified way. This architecture gives Hidream "Any to Any" capability: arbitrary modality input can support arbitrary modality output.

Tao Mei summarized Hidream's technical conviction in one sentence:

Image is the starting point to the world, and the gateway to reshaping it. In his view, image isn't merely image — it carries time and space: add temporal dimension to image sequences, and you get video; add stereo geometry, and you construct 3D worlds; complete interaction and operation with the physical world, and that's embodied intelligence. Behind this path is Hidream's complete vision of "from multimodal to native omnimodal, then to world model."

At the end of his WAIC speech, he summarized what this route means for AI's future in one sentence:

Large models give AI the power to understand and perceive the world; agents give AI the power to execute and act; native omnimodal world models close the loop of cognition, action, and feedback, pushing AI from content generation toward changing the physical world. According to official Hidream.ai disclosures, its self-developed UiT (Unified Transformer) native omnimodal architecture has surpassed 200 billion parameters in its closed-source Pro version; vivago R1, powered by HD-AgentOS full-process scheduling and governance, achieves an 85% content effective-usability success rate.

In June this year, Hidream.ai's commercial image model HiDream-O1-Image-1.5 ranked second globally on independent AI model evaluation platform Artificial Analysis with a 1265 ELO score, trailing only OpenAI and surpassing comparable products from Google, NVIDIA, and ByteDance.

From 2017, when Tao Mei led his team at Microsoft Research Asia to publish the world's first text-to-video paper, to 2026, when vivago R1 enables ordinary users to generate coherent long videos with one click. The evolution of AI video is happening faster than anyone anticipated.

When a short clip no longer needs a team of dozens, when a 120-minute series can be compressed to one month or faster, what new narratives will creators — previously shut out by cost and technical barriers — bring into being?

The answer is only beginning to be written.


Crossing is seeking independent contributors to write AI product and model reviews.

If you've written pieces like: "Hands-on with PixVerse C1[1]," "Hands-on with LibTV[2]," please contact zeo0811@gmail.com. Your email should include: ① personal introduction, ② AI review articles you've written.

We offer competitive rates. Looking forward to observing and documenting the AI era together 🎪