If you still don't get why Seedance 2.0 blew up, we made eight videos to show you.
One image, nine references — and Skylark cut a complete short film.
One image, nine references, and Skylark stitched together a complete short film.
👦🏻 Author: Jingshan
🥷 Editor: Koji
🧑🎨 Layout: NCon
For the past couple of days, ByteDance's next-gen video model Seedance 2.0 has been dominating the AI video space — just two months after its predecessor dropped.
Across social media and creator communities at home and abroad, users have been frantically sharing and resharing their generations. Every other post is racking up likes and reposts, and the comment sections all sound the same: "This time, the AI video model actually hits different."

Zoom out to early 2026, and this "hits different" quality becomes easy to explain:
AI video models are starting to understand the relationship between shots. In plain terms, they're learning to tell stories.
Seedance 2.0 is, in a sense, the signal that this turning point has arrived.
ByteDance's AI video creation agent Skylark has just rolled out this model. We spent a full day putting it through deep testing, verifying every viral trick one by one to see how far it can actually go.
Here's what we found.
Ancient-Style Roast Duck Explainer
A few keywords keep popping up in Seedance 2.0 discussions across platforms:
1) Reference capability is insane
2) Can precisely replicate subjects
3) Massive narrative improvement
4) High consistency with multiple subjects
We built a stack of test cases around these four points. Here's what we got.
First, the updated Skylark paired with Seedance 2.0 is especially good at "multiple reference subjects."
What this means: You can throw in a bunch of materials at once — several images, Douyin short video links, or video files. It references all of them together, and with simple, intuitive prompts, you can generate a video up to 15 seconds long.
Here's an example.
I came across a short video on Douyin recently. This niche is blowing up right now: history explainers with a background image, delivered in a weird, comedic tone and rhythm.

Meanwhile, I'd been seeing people make these "roast duck wrap explosion diagrams" — deconstructing a wrap layer by layer, laying out each ingredient separately.
So I had an idea: use an ancient-style character image as the host, and have Skylark reference the music and rhythm from that short video.
The whole thing would focus on just this one roast duck wrap image, explaining step by step how the pieces come together. Clear, rhythmic, satisfying.


Writing prompts in Skylark doesn't need to be complicated. Just follow your idea, making sure to clarify what you want, what to reference, and who's doing the talking.
I want to make an explainer video in the style of [reference video], abstract and fun, with [image 2] as the content, explaining how the roast duck wrap is assembled. Replace the host with [image 1]
Generation speed was decent. A few things about the final result are worth highlighting.
First, the shot transitions are smooth, character movements are stable, and the overall rhythm matches the music — no heavy "AI sheen."
Look closely at the close-up where the character rolls the duck wrap, and you'll notice very realistic focus pulls. The lens starts sharp on the host's lifelike face, then racks focus to the roast duck wrap as the camera pushes in.
Throughout this focus shift, visual element consistency holds up well.


Indian Comedy Sketch, Anime-Style Remake
Beyond full anime aesthetics, it can also convert real-person comedy shorts directly into anime versions — with closely matched movements and precise subject replication.
This is another thing people have been going wild with online. I've been seeing tons of Indian real-person comedy sketches lately:

I had Skylark make a similar version:
Swap in other characters, using two images I uploaded
The reference images were deliberately "crossover" ancient-style figures:


Here's the result. The two characters' movements connect without much drifting, the Indian comedy sketch replication is on point, and the characters' expressions read as more emotionally grounded.
"Corporate Drone" Extreme Sports Montage
After testing the examples above, Seedance 2.0's improvements in reference handling and subject fidelity felt pretty clear.
Beyond that, its narrative capability kept showing up in creators' content too. This mainly manifests in Seedance 2.0's ability to understand "images and video," connecting each shot to the next.
You're probably familiar with those extreme sports compilation videos on short-form platforms — rapid cuts, tight rhythm.
I used Skylark's interface: uploaded 8 images, then pasted in a Douyin short video link as reference. That's 9 reference materials total.
The prompt was simple:
Follow the camera movement and rhythm of [reference video], make me the same kind of short video but swap the extreme sports with my 8 uploaded images. Multiple shots, multiple angles, funny and abstract

The final result felt genuinely like a "short video."
The first frame grabs you, with a fisheye lens effect. Shot transitions between images are smoother. The music doesn't slavishly copy the original video; instead, it composes something new that follows the image changes, sounding more cohesive.
These creator shorts usually also have "artistic text overlays," with different text for each shot. Doing this manually is tedious, but Skylark can now add text throughout the entire video automatically. You don't need to frame-by-frame specify "write this here" — it reads the visuals and fills in appropriately.
Prompt:
Add artistic text introductions to the reference video, with a different narrative description for every shot. Don't make it too neat — give it an explosive feel
Look closely and you'll see: it basically follows the original video's rhythm, adding text on top, making the whole thing feel more polished.
Transitions are smoother too. Instead of "hard-cutting to the next image," it often gives a close-up first — say, zooming in on a crumpled paper ball in someone's hand from one image, then cutting to the next frame.

Multi-Subject Concert
From these examples, it's clear that "keeping characters from morphing" works much better now, and its video understanding has gotten stronger too.
To push this consistency capability further, I uploaded a photo from Koji's podcast, then found a short video link of a Queen stadium singalong — mainly to test how far its "multi-subject consistency" has come.
P.S. I noticed a couple days ago that you can no longer upload "real-person materials + copyrighted content," probably because Seedance 2.0's output is now so realistic. For the actual result, check out the demo below.


My concept: replace the concert's lead with Koji's character, plus some hip-hop flavor. Three main requirements:
-
Use Koji from the reference image as the character
-
Follow the reference video's shots and overall feel
-
Keep the music going, but add some Chinese hip-hop vibes
Here's Skylark's result. Even as shots cut, backgrounds shift, and crowds move, the lead's face and vibe hold steady — no wandering off.
I also tried "adding another subject" — say, an EVA Unit-01 with strong visual contrast:

The result: despite one being a real person and the other a mecha crossover, the overall blend is remarkably coherent.
Fisheye Lens Action Cam
For a closer look at real-world scene consistency, a "real person hands off action cam, recorded with fisheye lens effect" setup feels more "practical."
Note: the action cam in the video below was actually generated with AI image generation. You can observe the consistency closely:

Koji holds the action cam. The first half has almost no music, then suddenly powers on and switches to fisheye — the music hits instantly. The transition feels natural, and it's exactly the rhythm trick trending in shorts right now.
You can even get Seedance 2.0 to produce action cam selfie fisheye effects. Under this lens, Koji's facial details barely distort — consistency is remarkably high:

Solvay Conference Dances to Viral Choreo
Lately Douyin has been flooded with "make people in videos AI-dance" shorts. Like this one: the Solvay Conference photo dancing to "Green Apple Paradise."

On short-form platforms, you can always just upload an image and use effects to copy the trend. But if you want to add more characters to that image, or make a more complex version, that gets tricky — and that's where Skylark shines.
I had four characters I wanted to drop into the Solvay Conference photo, doing the same choreography as the people behind them, while swapping the music.

Here's the result. Overall it's stable, movements track with the music, and the effect is pretty addictive.
One detail I think is crucial: behind the young woman in the front row, second from left, is Marie Curie. This kind of "two people stacked front-to-back" is where AI dancing usually falls apart — movements easily get garbled.
But in this video, even Curie in the back keeps time with the music, and her movements don't glitch during transitions. The whole thing reads as solid.

To sum up: running these real cases, you can directly feel several practical capabilities it now has.
First, multi-material reference: whether multiple images or short video links, you can upload them all together — no need to lock to a single reference source. Even 8 images + 1 video reference works. And cross-modal reference is possible: video + music + images + text combinations.
Then, multi-shot continuous generation: a single video can break itself into multiple shots to tell the story, with shots that mostly connect rather than feeling like isolated fragments.
Also character stability — the same person doesn't randomly change faces or outfits across shots — and movement tracking: continuous actions like rolling wraps, chewing, dancing all read as smooth.
Its rhythm and camera movement mimicry is on point too, following reference videos' cut speed and motion patterns.
Auto-generated text and image captions come out directly, no frame-by-frame labeling needed. Finally, multi-person same-frame and front-back occlusion in more complex compositions mostly hold up without the whole frame collapsing.
In short, Seedance 2.0's viral moment reflects an industry threshold being crossed — and this time, the hype is deserved.
🚥
Seedance 2.0 is now fully rolled out. Come try it on Skylark ~~

