Who Is Mass-Producing the AI Short Dramas You Keep Scrolling Past? | Crossing's Hands-On Test

The model sets the ceiling; the agent sets the floor.

The model sets the ceiling; the Agent sets the floor.

👦🏻 Author: GaKi

🥷 Editor: Koji

🧑‍🎨 Layout: NCon

At the 2026 WAIC that just wrapped up, Hidream.ai officially launched its content creation agent — vivago R1. The product went live globally yesterday, with a fully upgraded domestic version called "Gouda" (够搭), positioned around the stable generation of complete 3-to-5-minute videos.

The launch lands right as AI video technology is advancing rapidly.

Judging by the latest Artificial Analysis blind-test rankings, competition among foundation models is fierce. Major tech companies at home and abroad, alongside AI unicorns, have all jumped in, each showing off strengths in visual fidelity, motion amplitude, and generation efficiency.

Foundation models keep getting better, but there's still a break in the workflow between what a model renders and a finished film — and that's exactly where Agent products add value, and where vivago R1 positions itself.

🚥

The Crossing team ran an in-depth hands-on test right away. Here's what the experience was actually like.

Hands-on: a 5-minute viral "Celestial Palace" feature, plus 6 TikTok e-commerce short videos

First, here's the access link for vivago R1:

https://goudaai.com/home. Once you're on the site, just type in a prompt or call up a Skill, use @ to add assets, and you can start creating.

If you've been spending time on social and short-video platforms lately, you've definitely noticed that short videos like the one below — celestial palaces, Chinese fantasy aesthetics — are everywhere. They've practically become a mainstream subgenre of AI short dramas. The core ingredients: wide establishing shots and colossal-scale aesthetics.

Now vivago R1 can directly output a roughly 5-minute celestial palace AI short drama in one go. Add a catchy, internet-native soundtrack at the end, and the vibe is just right.

To start, you can talk to vivago R1 in plain natural language — no node-building required. I first had it generate three 21:9 cinematic-format Eastern mythological epic scene images.

After that, a brief conversation with the vivago R1 Agent is all it takes: ask it to design a storyboard script referencing those images, and it will fully automate the visual design and asset generation — parsing the entire storyboard and establishing a coherent visual style.

Once the overall structure is locked in, it automatically moves to the next step, generating image assets in serial batches. Every scene, character, and prop required by the storyboard gets turned into an image asset.

One nice touch: when generating characters or props, it defaults to producing front, back, and side views of a character, or multi-angle views of an object — these serve as style anchors for the video generation that follows.

It then writes its own prompts for every storyboard reference image and every AI video clip to be generated, builds a to-do list on its own, and confirms clip by clip — there can be a dozen-plus verification stages in total.

This verification step is meticulous, and for good reason: without it, the final video could easily come out entirely unusable.

For example, vivago R1 verifies the storyboard length of every clip — clips aren't fixed at 30 seconds; they also drop to 15 or 19 seconds, staying consistent with the overall storyboard plan.

It then checks whether all reference images and asset references match up, whether every asset is covered, and whether any prompt has been abnormally truncated.

When a prompt is too long, the AI may silently truncate it — but the product's front end won't show that. A user writes a long, detailed prompt and then has no idea why the output looks so bad.

So this kind of careful internal checking is genuinely necessary.

At this point the full piece had 17 storyboards, mapped to 17 scenes and merged into multiple clips. All image assets were done, and the themes for the 7 video clips were written. Only at this stage does it ask the user to confirm whether to generate the video.

From the initial natural-language prompt all the way to this point, I barely interacted with the vivago R1 Agent at all — the Agent executed everything autonomously.

Next, the Agent generates the video clips one by one, following its overall plan, the image assets, and the prompts.

Each clip gets added to the canvas, and then everything is automatically edited together into one complete film.

Before showing the full five-minute film, here's a finished segment with short-video music added — the moment the music kicks in, the whole vibe clicks (ps. the music BGM was added manually in CapCut, licensed from "Bishang Guan" (《壁上观》), originally sung by Xiaohan Zhang; all other sound effects and BGM were generated by vivago R1).

The first version already came out remarkably complete — a usable 5-minute film straight away. I wanted to add more plot, so I generated it again.

I lightly edited the two videos and combined them into a roughly 6-minute mini AI short drama, Celestial Palace: The Returning Lamp.

A traveler in white, carrying an unlit lantern, spends a day ascending to the celestial palace to deliver the lamp to the moon. That's the whole plot. The film uses a variety of cinematographic techniques to showcase the architecture and style of the palace.

Overall, the plot hangs together tightly, and character consistency and style-building are solid — impressive given that the only style references for the entire pipeline were those first few images. Every other image asset was generated later by the Agent.

The final film has a genuinely poetic quality.

While testing this first celestial palace drama, I noticed vivago R1 has a fairly complete Skill Hub, with Skills for TVC ads, script-to-video, image editing, and more.

Some of the TVC ad Skills are especially well-suited for TikTok and Reels. There's one called Story AD Director, which I used to have the Agent produce six TikTok-style narrative product ads in one shot.

This Skill follows a fairly fixed structure — for example, dramatic conflict appears within the first 3 seconds, and the physical product gets placed naturally at key plot turning points. The overall framework is well thought out.

To use it, you click "Try Now," enter a prompt, and it automatically applies the Skill's structure to optimize your prompt.

My scenario: produce a full set of TikTok and Reels conversion ads with a live-action feel for a men's skincare face cream. The lead is a Black man who uses the cream across multiple settings — a wedding, an airport, a red-eye flight.

Once the Skill kicks in, vivago R1's Agent directly creates the lead's character design — generating, within a single image, multiple facial expressions, outfits matched to each scene, product display shots of the cream — and even draws out storyboards for the entire scene.

Below is a finished 30-second spot from one of those scenes. vivago R1 can directly output this kind of TVC-style short video.

This 30-second spot follows the classic skincare-ad narrative arc: the opening shows the lead's exhaustion and rough skin after a red-eye flight; the middle presents an instant-repair sequence of washing his face and applying the cream. The ending uses a cold-to-warm color-grade shift to reinforce the transformation, then wraps up cleanly with product imagery and a brand tagline.

This Skill makes producing this kind of TVC short video genuinely convenient. On platforms like TikTok, 30 seconds is already on the longer side for an ad, and it can output many spots in one batch with high consistency across videos — including character and product consistency.

We happened to hand-build a short-video stitching tool with GPT 6 Astra, which makes it easy to show the final result.

Here's the final assembled video:

Look closely and you'll notice something striking: although the six 30-second TikTok TVC ads are set in completely different scenes, their overall structure is remarkably consistent. Around the 3-second mark, for instance, the lead always appears frowning — in a negative state, facing some problem in his life.

Then, all six ads introduce the product at almost the exact same moment. At the 17-second mark, for example, the lead actually applies the men's face cream.

All in all, even though I only used very simple natural-language prompts, once paired with this Skill, the vivago R1 Agent lays out different scenes based on the Skill and my prompts — while maintaining the same narrative structure across all of them.

This matters a lot when batch-producing TVC ads: you need consistency, but you also need to avoid prompts drifting so far that the visuals fall apart.

So keep the scenes consistent, keep the narrative structure consistent — hook first, then solve the problem, then the product makes its positive entrance. That's very much in line with how TVC ads work.

The product logic of vivago R1

After hands-on use, vivago R1's product thinking comes down to two layers.

【1】Interaction: pure conversational mode

The entire creation process requires no node-building and no manual parameter settings for image models or video duration. vivago R1 encourages users to create by talking to the Agent in natural language.

In practice, it leans heavily on Agent interaction — nearly the entire pipeline, from script breakdown to shot planning, role assignment, and pacing, is handled by the Agent, which delivers the finished film.

The team uses an interesting analogy: "the entry-level DSLR of AI video creation" — simple enough for non-professionals to pick up, stable enough for professional creators to deliver work.

【2】Long-task orchestration

Videos longer than 5 minutes count as long tasks. No model today can generate several minutes of coherent video in one pass, so an Agent is needed to orchestrate the whole creation process.

R1 hands the orchestration to the Agent, which plans the storyline from the prompt and keeps characters, products, and scene styles consistent across shots. Users can also revise any individual shot through conversation.

For an Agent to own the entire creation flow and minimize manual work from users, the product has to smooth out the whole pipeline.

R1's ability to connect this pipeline is tied to the fact that the team behind it — Hidream.ai — started from foundational multimodal model R&D. Whether you can run this kind of long pipeline end to end depends heavily on how deeply the team understands model capabilities.

Their earlier HiDream-O1-Image reached the top ranks of the Artificial Analysis image leaderboard, and their interactive world model HiDream-O1-World scored 80.9 on WBench.

So it stands to reason: only a company that has wrestled with the hard problems of model development itself knows how an Agent product should lower the barrier for users. To spare users the grind of repeatedly prompting video clips and assembling them into a film, the company needs to know exactly which steps require engineering logic as a safety net.

🚥

"The model sets the ceiling; the Agent sets the floor."

That line captures the 2026 landscape of AI video models and their commercialized Agent products rather well.

If Agents' orchestration and comprehension abilities keep evolving, the boundaries of AI video creation will very likely expand further. As commercial potential gets unlocked, AI video creation will become a genuinely lucrative business — and those commercial returns will feed back into better AI models.

So we're eager to see just how big a commercial impact AI video models can make this year, and how far Agent products can go.