After AI Comic Dramas Hit 2.5 Billion Views, the First Batch of Vertical Models Emerges | Hands-On With PixVerse C1
What does a vertical model purpose-built for short dramas, anime, and film and TV content creation actually look like?
What does a vertical model trained specifically for short dramas, anime, and film/TV content creation look like?

👦🏻 Author: Jingshan
🥷 Editor: Koji
🧑🎨 Layout: NCon

The 2026 Spring Festival box office brought a striking number: according to Monnfox stats, AI-generated anime dramas surpassed 2.5 billion views, capturing nearly 30% of the short drama market.
Just a year ago, AI anime dramas were still a novelty. Most people who stumbled across them found the visuals slightly off, the movements stiff, the whole thing uninteresting — they'd swipe away after two seconds.
But this year, the entire AI anime drama market has seen major growth in scale, daily new releases, and user base. The question is no longer whether it can be done.
Which raises a very practical problem.
General-purpose video models keep getting stronger, but when applied to the specific contexts of anime dramas and short dramas, they still have clear weaknesses. These aren't problems you can solve by simply scaling up the general model further. They exist because general models were never optimized for these scenarios during training.
It's against this backdrop that on April 8, 2026, AISphere (PixVerse) released C1, a model with a sharply defined positioning: the world's first large model for the film and television industry, a vertical model trained specifically for short drama, anime, and film/TV content creation.
🚥
Let's start with PixVerse's recent flurry of model releases, talk about where C1 sits in their product lineup, and then share our hands-on tests.
PixVerse has been releasing a lot of models these past six months
If you follow the AI video space, you've probably noticed that PixVerse's release cadence has been almost dizzyingly dense since the second half of 2025.
Here's a quick rundown.
In H2 2025, the V5 series went through several iterations in quick succession. V5 itself was a major upgrade to general capabilities — motion effects, high-res image quality, character consistency, and prompt adherence all got pulled up. We covered it in depth in "PixVerse V5 Drops Surprise Launch, We Tested It Deep on 'Pai Wo AI' First". V5 Fast pushed 1080P video generation under 30 seconds.
By December, in our article "Hands-on with Pai Wo AI V5.5: AI Video Creation No Longer Needs Complex 'Workflows'", we found that V5.5 did something fairly crucial — it was the first to support one-click generation of storyboards plus audio, essentially moving from "generate a video" to "tell a story."
January 2026 brought R1. Positioned as the world's first general-purpose real-time world model, it basically transformed video generation from "wait for results" to "real-time interaction" — users could change directions while the scene was still generating. This was completely different logic from the V series, pursuing an interactive experience.

In February, V5.6 ranked #2 globally on both the i2v and t2v leaderboards from Artificial Analysis, second only to Seedance 2.0.
Then on March 30 came the V6 model.
V6 was another major upgrade to the general-purpose flagship — character realism, complex motion, physics simulation, and audio-video sync all got maxed out, with maximum generation length extended to 15 seconds.
C1, which followed, went in a different direction entirely.
It has no derivative relationship with V6, nor was it fine-tuned on top of V6. C1 is a vertical industry large model independently trained by AISphere, with specialized optimization for fight choreography, spell/technique special effects, multi-panel storyboarding, and other such scenarios.
This also marks the first time AISphere has trained a dedicated model for a specific industry.
Hands-on with C1
Below are several tests we ran with C1, covering its core capabilities. Each case includes the actual prompt we used.
Case 1|Multi-panel storyboard direct output
Multi-panel storyboard direct output is one of C1's headline capabilities this time around.
You can simply feed it a nine-panel AI anime drama image. The model reads the visual relationships on its own, understands what each panel is depicting, then strings these shots together into a complete video automatically. You basically don't need to manually break down shots or fill in narrative logic anymore.
The operation is straightforward too. Just select the C1 model in the input box. Maximum resolution goes up to 1080P, maximum duration 15 seconds. Audio is generated together, no extra processing needed. It's essentially a complete image-to-video direct pipeline.

I tried a nine-panel storyboard in ancient fantasy style, with each panel respectively: distant mountain gate, two figures facing off, drawing swords and striking, sword qi collision, shockwave spreading, characters recoiling, sky splitting open, light pillar shooting upward, final freeze frame.
Oh, and this entire set of images can also be generated directly in PixVerse, which likewise offers the Nano Banana Pro model.

Then you can pass this image directly to C1 as reference, with the following prompt:
Ancient mountain sect entrance, two cultivators facing off, drawing swords against each other, true qi collision producing particle explosion, shockwave throws both sides back, a golden light pillar splits open, final freeze frame with silhouettes of figures against glowing sky. Dynamic camera, cinematic lighting, anime style.
The results were pretty solid. The transitions between panels flow smoothly, and character costumes and color schemes stay consistent across different shots.
The camera movement rhythm varies. The face-off uses a slow push-in, the collision moment accelerates, and the final light pillar gets a slowly pulling-back wide shot.
Several frames came out particularly well. For instance, at the moment of sword impact, the frame has a slight shake and light particles scatter outward — that particle feel common in anime dramas, fairly close to how Japanese and Chinese anime typically render it.

Another example is this spell technique effect, which stands out as one of C1's more obvious capabilities this time.
After the two clash, a light pillar directly tears through the sky. You can see debris getting kicked up, the frame shaking, and character footwork changing. Overall it hews much closer to anime drama visual logic.

Case 2|Fight choreography
Another noticeable thing about C1 this time: it's not just spell effects — fight choreography continuity is also more stable. Including basic physicality, plus special effects during combat, the overall package holds up.
For instance, I wrote a prompt leaning classical wuxia, not xianxia fantasy — a more grounded, realistic style.
The overall look is dark and cinematic. The setup: twin martial artists, brothers turned enemies, dueling on a rainy rooftop. Mainly sword techniques exchanging back and forth, with light particles, water droplets as details, plus changes in facial expression.
Prompt as follows:
Two twin martial artists, turned against each other, fighting fiercely on a rooftop in the rain. Fast-paced sword techniques, one rapidly slashing while the other parries and counters with a spinning kick. With each clash, rainwater splashes. Slow-motion captures the instant of blade impact, then returns to normal speed. Cinematic wuxia style, dim and melancholic lighting.
The spatial relationship between the two characters stays stable throughout, with no clipping. Raindrop splashes during combat look fairly natural.
There's a speed ramp from fast to slow. The frame of blade impact slows down, raindrops hang suspended, then speed returns to normal.
Several things stand out in this clip. Character expressions are more detailed, especially micro-expressions like furrowed brows. The effect of water droplets hitting faces is rendered. Combined with lighting and highlight shifts, the overall feel is more realistic.

The C1 model actually supports multiple anime drama formats, including 3D anime.
I modified the setup from the above prompt, no longer requiring twins but swapping in two martial artists with different appearances, same rooftop duel setting.
In 3D anime scenarios, facial expressions are actually harder to pull off than in 2D anime, but C1's performance here is decent.
Looking closely, you can see light particle and water splash details, especially the effect of droplets hitting character hair. Going further down, even the way the two grip their swords shows differentiation — different hilt designs, visible veins on hands.

And looking even closer, you'll find action details rendered too.
For instance, when the white-clad martial artist executes a back kick, it kicks up a layer of mist. Later when both move across the rooftop, almost every step stirs up mist around their feet.

Case 3|Spell/technique special effects
Spell special effects in short dramas are a major draw in AI combat anime dramas right now. As seen above, C1 has been strengthened in this area, so this time we can specifically examine how it handles high-intensity lighting and more magical spell effects.
For instance, I wrote a scenario: three cultivators in a burning temple, fighting a massive flame demon. This scene needs an ice shield technique.
C1's "understanding" and reasoning of such spells holds up — like having one cultivator deploy an ice shield, then stacking lightning effects to collide with the demon's flames.
Prompt as follows:
Three cultivators fighting a massive flame demon in a burning temple. One cultivator deploys an ice shield, the green-robed cultivator wields a glowing spear and lunges at the demon while lightning strikes from the sky. The demon swings its giant claw, shattering the ice shield. Fragments scatter, flame particles float in the air. Epic anime battle, dynamic camera angles.
The resulting video:
Looking closely, while there are still some minor flaws, the spell effects overall work. The shattering sensation when ice shield meets demon, the flame trail when the green-robed cultivator swings his spear, and the environmental changes triggered by lightning from above — all come through well.
Especially the segment where the cultivator releases lightning mid-air — it looks quite impactful overall.

Video post-processing: Edi
Of course, the video above still has some minor flaws. These are common issues when using C1 for this type of AI anime drama. And they're hard to fix in post-production through editing, since the footage itself is generated.
However, PixVerse itself has a companion capability called Edi, essentially a video post-processing tool.
Simply put: you generate a video that's roughly 60-70% there overall, with some minor issues in places — maybe an element that shouldn't be there, or you want to swap out a prop in a character's hand. Instead of regenerating from scratch, you use Edi to modify directly on the original video.
The industry logic behind C1
Pulling back from these hands-on tests, there are a few things worth discussing.
[1] First, the emergence of vertical models signals that the AI video赛道 is beginning to "specialize."
For the past two years, all AI video models have been digging deep on generality. This made sense early on — foundational capabilities needed to reach a certain baseline first. But once leading general models all hit 90+ scores, squeezing out those remaining 10 points became increasingly difficult.
And every specific industry has its own "remaining 10 points."
Anime dramas need fight fluidity, special effects realism, and storyboard consistency — these don't completely overlap with what general models pursue as "visual beauty" and "natural motion." C1's emergence is, in a sense, saying: after general models hit 90, the next battleground may be in vertical scenarios.
[2] Second, AI anime drama production is becoming an "industrial pipeline."
For many practitioners in the AI anime drama space, the most deeply felt change is probably in workflow. Previously, making one episode of AI anime drama meant generating videos one by one, then checking, selecting, and editing them one by one — an extremely fragmented process.
C1's multi-panel storyboard essentially merges "video generation" and "storyboard logic" into the same step. Combined with Edi's post-processing capabilities and so on, a toolchain spanning generation to modification to batch deployment is taking shape.
For teams pushing dozens or even hundreds of short dramas daily, the value of this toolchain far exceeds any single capability improvement.
As a side note, C1 currently has a limited-time free experience window: April 7 to April 17, with full C1 functionality available for free on web clients both domestically and internationally. API customers can also get 20% off within 7 days of launch. If you're interested, this window is worth trying.
Back to what we were discussing.
The AI anime drama赛道 isn't lacking for heat anymore. Attention has shifted from whether AI can make anime dramas.
The real question is: who can push this赛道's production efficiency to the next level?
C1's answer is straightforward: train models specifically for industry scenarios, tackling the most painful points in fights, special effects, and storyboarding one by one, then pair them with post-processing tools and API interfaces to form a complete toolchain.
How far this path can go ultimately depends on its performance in real production environments. But at least from what we can see so far, AI anime drama-specific models do have things that others can't do in the anime drama/short drama赛道.
The race for AI anime dramas may have just entered its second half.

