Over the course of a five-day launch event, Vidu unveiled a complete content production pipeline.
In five days, Vidu delivered a complete production toolkit.
**
5 Days, Vidu Delivered a Complete Production Toolkit

👦🏻 Author: Jingshan
🥷 Editor: Koji
🧑🎨 Layout: NCon

The team at Vidu Shengshu Technology has definitely had a busy few days.
From January 27 to 31, they held a release every single day for five straight days: Q2 Reference-to-Video Pro, Agent 1.0, Subject Community, Q3, and an ecosystem plan — all launched back-to-back.
This cadence is unusual in the AI industry. Most companies spend weeks building hype for a major product drop, then take more time iterating after the launch event.
The goal of this round of updates is clear: pushing AI from "capable of generating video" to "entering the content production chain."
This also follows the same path they've emphasized from the start: "production-grade AI video capabilities."
To understand this release, look at a concrete scenario. Say a short drama team wants to use AI to create a 16-second story clip. What do they need?
Controllable character expressions and movements. Replicable and editable special effects. Dialogue that matches lip movements. Camera cuts with narrative rhythm. Background music and sound effects that fit, plus the ability to quickly produce in bulk.
Put these requirements together and you already have a complete production pipeline. As it happens, Vidu's five days of announcements basically fill in the gaps at each of these stages.
🚥
If you view these five days of releases as a puzzle, you'll find they form a complete production toolkit. What follows is our breakdown of these five "products" across five days, seen through the lens of the production workflow.
What Did Vidu Do to Complete a "Full Narrative Production Pipeline"?
From the perspective of real production workflows, AI video models have evolved rapidly over the past two years. Each iteration has pushed the entire process forward — from motion graphics to long-form video, from blurry to high-definition, from 5 seconds to 10 seconds in length, with near-monthly improvements.
But these advances have been more like icing on the cake. It's been hard to lift the entire professional creative process up another level. Vidu's releases this time are trying to solve three problems: how to make generated results controllable, how to make content tell complete stories, and how to optimize workflows for specific scenarios.
So Vidu released two models, one Agent, one subject community, and one ecosystem plan over these five days.
Overall, the main goals fall into three paradigms:
[1] AI video generation controllability will improve, attempting to break free from "random generation" — this corresponds to Q2 Reference-to-Video Pro;
[2] In real production scenarios, models should be able to deliver "stories" — this corresponds to the Q3 model;
[3] For vertical scenarios, there will be a solution: Vidu Agent 1.0.
Behind all this are two underlying ideas:
[1] Creativity is an asset, and AI video is becoming a new content ecosystem;
[2] Creators should be "one part of the ecosystem," no longer isolated.
Let's look at each layer.
First, Vidu released two models this time: Q2 Reference-to-Video Pro and Q3. Though the model numbers keep climbing, these two aren't simply "Q3 is greater than Q2 Reference-to-Video Pro." Q2 Reference-to-Video Pro leans more toward "reference," while Q3 is Vidu's next-generation audio-visual synchronous generation model.
Let's start with Q2 Reference-to-Video Pro.
1) How Far Did Q2 Reference-to-Video Pro Push "Controllability"?
For AI video to truly be used in production, one problem must be solved: how to make generation results shift from "random surprises" to "controllable expectations." This release of Q2 Reference-to-Video Pro pushes that capability forward.
Launched on January 27, Q2 Reference-to-Video Pro is billed as the world's first "everything is referenceable" video model. This isn't just about expanding the reference scope.
Past AI video models could mostly only reference static elements in images: character appearance, scene composition. But in real creative scenarios, creators often need to reference dynamic, abstract elements. The logic behind how an effect is achieved. The expression of an emotion. The texture of a material. The rhythm of a movement.
But if you wanted AI to reference these, that was basically pretty difficult.
Where's the problem? The reference scope was too narrow.
Models only recognized "this person" or "this scene," but couldn't understand "how this person's transformation effect works," "what emotion is behind this expression," or "where the texture details of this material lie."
What does "everything is referenceable" mean?
Simply put, basically anything visible in a video can be used as reference. Special effects, emotions, materials, movements, and so on. Q2 Reference-to-Video Pro supports multi-video, multi-image references, allowing users to combine different reference materials to create new effects.
For example, making the character from Image 1 have the special effect from Video 1, while referencing the movement from Video 2, plus the scene atmosphere from Image 2. This composite reference capability dramatically increases creative freedom.
Using Q2 Reference-to-Video Pro's intuitive prompt design, you can make multi-subject reference videos.
For instance, I had Q2 Reference-to-Video Pro directly reference the subject of this image and the effect of this video:

And it produced this kind of "reference-to-video" result:
Reference is only the first step. Another dimension of controllability is editability.
Q2 Reference-to-Video Pro's new video editing features let creators perform fine-grained add, delete, and modify operations on generated content. More importantly, even after multiple edits, consistency of the main subject in the frame can still be maintained. This is crucial in actual production, because no model can achieve "perfect on the first try."
Even with repeated edits, "consistency" of visual elements can be preserved as much as possible. This matters in real-world creation, because creators often need to iterate repeatedly until satisfied.
Put simply, AI video models aren't good enough to be "no-edit" yet. Creators need a "shovel."
If every edit broke consistency, this tool wouldn't be usable. Q2 Reference-to-Video Pro's consistency preservation lets you modify step by step, changing one thing at a time, until the entire frame reaches what you want.
The barrier for special effects has been dramatically lowered.
2) Q3 Pushes the "Narrative Problem" Further
While Q2 Reference-to-Video Pro solves "how to control the image," Vidu Q3 — released on January 30 — answers: how to make AI generate story units with complete narrative capability.

A complete story clip needs more than just images. Dialogue, sound effects, background music, and camera cuts all need to work together to tell a story.
Vidu Q3 is the world's first model to support 16-second direct audio-video output. The key words here are "direct" and "16 seconds."
"Direct" means sound and image are generated together, synchronized at the model level.
The 16-second breakthrough isn't just a duration number. Traditional AI video workflows are "generate image first, add sound in post" — sound and image are produced in two separate stages, with coordination dependent on manual adjustment. Q3's "audio-visual synchronous generation" capability lets sound and image achieve native synchronization at the model level.
Like this piece, a story short directly generated by Q3, using 4 camera cuts to tell a relatively complete plot:
Beyond audio-visual sync, Q3 also supports automatic multi-camera switching, narration, two-person dialogue, sound effects, music, and more.
Looking at benchmarks, Q3 ranked first in China and second globally on the latest list from international AI benchmarking organization Artificial Analysis, surpassing Runway Gen-4.5, Google Veo 3.1, and OpenAI Sora.

This combination of "audio-visual sync + multi-camera switching" lets Q3 attempt to carry the full arc of a complete shot sequence — exposition, rising action, climax, resolution — achieving fluidity of rhythm, emotion, and narrative, fused into a unified whole.
Thus, Vidu positions Q3 as a next-generation video model "born for drama." It upgrades AI video from "asset generator" to "story segment producer," embedding directly into the front end of screenwriting and storyboarding creation.
Beyond these two most closely watched AI video models, running through all the above features is a module called "Subject Community."
3) Subject Community Lets Professional Expertise Circulate Too
When model capabilities achieve technical breakthroughs, where's the next bottleneck? In the reuse and circulation of creative experience.
Even the best model, if every creation starts from zero — parameters, prompt tuning, and so on — production efficiency is still limited by individual experience accumulation.
Imagine if there were a resource library with ready-made camera movements, effects, compositions, and atmospheres that you could just @ and call up directly. What would that be like?
This is what Vidu's Subject Community, released on January 29, is doing.
It's the world's first AI video creation subject community, containing 8 types: camera movement, composition, narrative, style, scene, performance, moves, and atmosphere. There are currently 200+ subjects, with official commitments to ongoing updates.
What's a subject? It can be understood as a reusable creative resource. A shot language, an effect style, an emotional atmosphere, a set of performance logic — these professional capabilities that originally required long-term practice to master are now packaged as directly callable "assets."
For example, "tense, oppressive atmosphere" as an abstract concept is one subject. Call it up, and your video takes on that tension, with color tone, lighting, and rhythm all cooperating.

Subjects can be freely combined.
You can simultaneously call @camera movement + @atmosphere + @performance, stacking multiple subjects for more complex effects. The workflow after adding the "Subject" module is: when selecting "reference objects," go directly to the "Subject Community" to find them.

"Subjects" that creators have previously designed, or that other creators have made, can both be uploaded to this community.

Another breakthrough of Subject Community is its ecosystem mechanism for sharing, trading, and interaction.
Professional creators can upload their polished shot languages and unique effects as subjects for others to use for a fee. Ordinary creators can also directly build on others' subjects for secondary creation, iterating through collaboration.
For ordinary creators, Subject Community lowers the professional barrier.
You don't need to learn complex techniques — just know what effect you want, then @ and call up the corresponding subject. Creativity is no longer limited by individual experience accumulation, but can continuously appreciate in value through circulation, becoming a kind of "compound interest" asset.
4) Vidu Agent 1.0 Wants to Reshape the "Marketing Video Production Chain"
Technology is only truly useful if it can be applied to real business. For advertising brands, what they want isn't flashy special effects demos, but video content that clearly explains product selling points, makes users remember, and gets them to buy.
Vidu Agent 1.0, released on January 28, is mainly solving this. Its core capability is one-click video generation. You just input requirements, and it automatically generates a complete video — similar to a general-purpose Agent effect.
Under the hood, Vidu Agent is actually a multi-Agent collaboration system that's already been refined considerably. You can think of it as: several AI Agents working together with divided responsibilities.
Currently Vidu Agent contains 7 specialized AI roles total. Someone's responsible for scripting, someone for shot breakdown, someone for organizing assets, someone for generating visuals, someone for voiceover and music, someone for editing rhythm, and someone for final quality control.
Each AI only does its own segment, then together they complete an advertising video.
For users, it's simple to use. You just write one requirement, provide one product image, and the entire process kicks off. Everything from script to visuals to sound and editing proceeds automatically. Along the way you can also manually adjust shots and structure to make the result better fit your goals.
The problem being solved here is: even if AI can produce good results, if it requires writing complex prompts and tuning parameters every time, the barrier is still too high for most creators.
Especially since traditional AI video marketing production workflows are actually quite long.
First figure out product selling points, then write the script, then do storyboarding, then round after round of prompt revision and asset generation, and finally post-production editing.
Each step basically requires professional oversight, so overall it's very time-consuming and labor-intensive.
Vidu Agent's overall experience this time is built for "marketing scenarios":

Upload an image, specify product selling points and concrete application scenarios, and the Agent automatically generates a storyboard script:

This Agent has more scenarios where it shines, with the most common being "product advertising." The Agent's approach: you just give it goals and assets, and it compresses much of that long middle chain for you.
This matters.
Because many creators can't, or don't know how to, "drag" subjects and write structured prompts like in the Q2 Reference-to-Video Pro and Q3 model workflows from the very beginning.
For marketing teams and self-media creators, Agent solves bulk production efficiency problems. Beyond efficiency, the more critical factors are cost reduction and room for trial and error. It's an Agent that understands products better — it can combine product characteristics to generate selling points with appeal and hit video structures, generating high-conversion advertising and marketing videos.
5) Yes, Vidu Global Ecosystem Plan: Get Tools Used, Keep Them Going
Who makes those professional subjects in Subject Community? Who uses the bulk-generated content from Agent? Who ultimately watches the short drama clips generated by Q3?
Tools, however good, if unused and unparticipated in, are hard to sustain.
The answer to all these questions points to the same thing: creators and users are the ones who make tools truly get used. Put simply: get more people using them, and let users earn returns.
On January 31 they released the Yes, Vidu Global Ecosystem Plan, directly putting up "100 million credits + tens of millions in prize money" to give resources to creators and partners, making everyone more willing to invest long-term. This mechanism is actually sincere — no convoluted twists and turns, with clear benefits.
For example, this "Artist Plan 2.0" below lets creators earn credits by generating quality content, sharing subjects, making tutorials, and so on. Credits can be used to generate more videos or exchange for memberships.

The certified instructor system follows similar logic. If you're skilled at using Vidu for a certain type of video and willing to share methods, you can become an instructor and get resource support. For enterprises there's also a partner system, with credit pools and business opportunity pools.
When creators can get returns, instructors can monetize expertise, and partners can continuously receive opportunities, the ecosystem will run on its own.
For tools to survive, the people using them must benefit first.
If you connect these five days of releases together, you'll find Vidu is trying to say something more practical: AI video creation in 2026 needs to achieve "capable of completing whole narrative segments," or attempt to handle the full video creation workflow.
This can be broken into three layers.
At the technology layer:
[1] Q2 Reference-to-Video Pro solves reference and consistency problems, letting creators control the image;
[2] Q3 brings audio and visual together into the model layer, letting segments tell stories;
At the application layer:
Subject Community turns shot languages and creative experience into reusable assets; Agent compresses complex workflows into a single requirement;
At the ecosystem layer:
Creators and partners participate together and share in returns.
The three layers together string the entire production pipeline together: from ideation → content creation → adjustment → asset reuse → publishing → monetization, there's basically a corresponding tool for each step.
Viewed individually, each product looks like just a feature upgrade. But used together along the production flow, they become a complete set of tools ready for real work.
Technical capabilities are more or less in place. What really matters going forward are some more practical questions, with three directions worth continuing to watch:
[1] Will a new class of creators emerge? Will there be a cohort of people who start making content with this toolset from day one?
[2] Will new content forms grow? 16-second complete stories, plus audio-visual synchronous generation — could this bring content formats different from current short videos?
[3] Will production capacity be amplified? For AI video content like short dramas and comic dramas, can output volume and production speed directly step up a level?
These changes will directly determine whether AI video stays in the demo phase, or enters large-scale commercial deployment.
🚥
Back to that short drama team from the beginning.
Now they want to make a 16-second AI short drama. Vidu's answer over these five days is they can do it like this:
Q2 Reference-to-Video Pro controls effects, emotions, and materials. Q3 generates 16-second audio-visually synchronized segments. Subject Community @-calls camera movements, atmospheres, and performances. Agent bulk-produces and rapidly iterates. Join the ecosystem plan for traffic and commercial partnerships.
Through one complete workflow, basically everything needed is there — each step has a corresponding tool, and each directly connects to the next. From this complete release, Vidu has already built out the tools, workflow, and creator system.
Next, it's up to creators to use this content toolchain to complete their own new narratives.

