PixVerse V5 Drops Surprise Release — We Put It Through Its Paces First on "Pai Wo AI"
After two years in the long game, AISphere is still sprinting ahead
**
Two Years In, AISphere Is Still Sprinting

👦🏻 Authors: Xiaoju, Jingshan
🥷 Editor: Koji
🧑🎨 Layout: NCon

Video has long since become the default language of our era.
From the endless scroll of social media to the grand visual narratives of the silver screen, we're all deeply immersed. And as AI technology has collapsed the barriers to video creation, a more level playing field — an "AI video model arena" — has opened for business.
The list of competitors is lengthening by the day: Google's Veo, Runway's Gen-3 have entered forcefully, while domestic newcomers like Hailuo AI, Keling AI, and Dreamina are "holding their own."
The pace of iteration across the entire industry can only be described as white-hot.
In this race, one homegrown "all-rounder" product has drawn particular attention — PixVerse (AISphere's domestic brand, "Paiwo AI"). Just recently, they announced their global user base has surpassed 100 million — no small feat. AI video is unleashing massive creative energy across the world at an astonishing speed.
Today (August 27), this user-heavy "all-rounder" completed another leg of the race, officially launching its eighth version in under two years — the PixVerse V5 self-developed video generation model. After less than 24 months of iteration, PixVerse V5 has delivered its latest report card to competitors, users, and the market.
Can this model of "fast" and "frequent" translate into "good" and "strong" results?
To answer this, we prepared 10 test cases covering both "model capabilities" and "various scenarios" for a comprehensive, in-depth evaluation of this AI video generation product. We tested the domestic version "Paiwo AI," which has been live for just over two months and remains synchronized with the international version in features, templates, and models.
Four-Dimensional Model Testing
First, let's look at the technical advances in PixVerse V5. I learned that this V5 brings substantial upgrades to core technologies like distillation, RLHF (reinforcement learning from human feedback), and unified feature spaces.
This PixVerse V5 self-developed video generation model, over two-plus years:
【1】In generation speed, maintains the rapid output from V4 — 5 seconds to generate shareable short clips.
【2】In aesthetics, more realistic, more natural, with smoother motion amplitude.
【3】In prompt comprehension, more precise than before.
PixVerse V5's Text-to-Video and Image-to-Video have ranked second and third respectively on Artificial Analysis — impressive results for a startup platform on this AI "college entrance exam."

Now, let's see the actual tests.
1) Consistency Preservation: Multi-Panel Oil Painting Through Four Seasons
First, consistency preservation. This requires PixVerse V5 to understand a slightly complex visual composition, then perform precise local transformations while maintaining "overall consistency."
The audio in the video was also automatically generated by PixVerse V5.
We uploaded a multi-panel painting containing 12 independent oil painting panels, each depicting a spring scene:

Then used a prompt of just seven characters:
Turn spring scenery into autumn scenery
This meant PixVerse V5 needed to recognize 12 independent yet interconnected spring scenes within the image, then process them sequentially into stylistically unified autumn landscapes.
The results:
You can see that in the first four seconds, PixVerse V5 applies some imagination to three main panels — for instance, in the top panel where a person walks, the path ahead keeps extending.
In the latter four seconds, the remaining nine panels shift from green to autumn yellows and reds all at once. Though somewhat abrupt, the overall consistency holds up reasonably well.
2) Dynamic Effects: Skiing and a Confident Smile from an Orange-Haired Girl
Next, testing "dynamic effects" — mainly whether it can render everyday life naturally and meticulously in motion.
PixVerse V5 upgrades motion trajectories and amplitude for realistic, stable performance. We generated a skiing video with solid results:
Then, I tried a different style, using a facial close-up to see how Pixverse V5 understands micro-expressions.
The prompt:
Dynamic angle, orange and white, grainy texture, skin texture, enhanced artistic atmosphere, green background, thick impasto + washed gouache, paint smearing + texture, sophisticated atmosphere, layering, visual tension from contradictory elements in organic fusion. Joy, warmth, girl smiles at camera.
The final result:

The girl is bright, confident. The orange-green pairing doesn't clash — it gives the image more vitality.
In terms of expression, PixVerse V5 chose a slightly retreating camera angle. As the camera pulls back, the girl slowly opens her eyes, corners of her mouth spreading to reveal cute canine teeth.
3) Visual Quality: Collage-Style Animation
Collage style frequently appears in brand advertising — Hermès, Nongfu Spring. I had PixVerse V5 generate a Hermès-inspired animation, with the prompt focused on scene description. With this, ordinary parents can make picture-book stories for their kids.
The prompt:
A young boy rides a galloping white horse through a fantastical forest, giant broad-leaf plants, blooming exotic flowers, a cheetah hidden behind leaves, a parrot perched in treetops, a rabbit in the underbrush, rendered in fine line engraving, color palette of black, red, orange and white, soft tones, finely detailed, Hermès style.
You can tell that visually, PixVerse V5 presents very fine lines, restrained but sophisticated colors, and an overall fairy-tale quality.
The forest scene also demonstrates fairly good spatial depth, with no major hallucinations.
The boy and white horse in the foreground, exotic flora in the midground, and the leopard concealed behind vegetation in the background — visual quality remains fairly refined across multiple spatial layers.
If you look closely at the textures of flowers and trees, the overall elements feel quite harmonious.
4) Prompt Adherence: A Girl Late for School
We used a text-to-video scenario to test PixVerse V5's ability to follow prompts.
The prompt:
An alley lined with paper mulberry trees, dense foliage forming a chaotic green canopy, self-built houses of varying heights and ages on both sides, concrete steps or wire mesh fences on balconies and doorways, morning glories, water spinach and garlic sprouts growing everywhere, some with folding chairs and plastic tables, bananas and oranges on the tables, a girl around thirteen rides past on a bicycle, backpack on, speeding by, seems late, somewhat flustered. Realistic and natural, clear brushstrokes, high saturation, light-shadow contrast, cinematic quality. Marker pen hand-drawn illustration style.
From the results, PixVerse V5 did generate the complete scene as requested: a leafy alley, the lived-in atmosphere of fruit on display, natural lighting and depth of field throughout the lane, matching everyday perception.
This shows its "language comprehension" is solid — it can convert text into concrete visual elements, and the camerawork tracks with character movement:
In terms of character performance, even without dialogue, viewers can tell this is a girl rushing to school. V5 conveys cycling speed through flying hair and fluttering sleeves.
If you look closely, you can even see the girl's brows furrowed, mouth slightly open — a "gonna be late" expression.
In style, PixVerse V5 does achieve a hand-drawn feel, with vivid colors and nice lighting effects, somewhat like watching an animated film.
Through these four tests across different technical dimensions, PixVerse V5 demonstrates fairly balanced comprehensive capabilities. Of course, there's still room for improvement in certain details and stylistic consistency.
Now, let's look at six specific scenario tests.
Six Specific Scenario Tests
1) Line-Art Figure Against Cluttered Background
This test primarily examines PixVerse V5's handling of character movement against "cluttered backgrounds," using a line-art comic sketch style with complex, graffiti-like elements.
The prompt:
A kung fu master in white practice clothes training and throwing punches. Line-art figure, full body, pencil sketch lines, graffiti, comic, chaotic.
The challenge: punching involves full-body muscle engagement, easily producing stiffness or mechanical repetition, and hallucinations frequently emerge when integrating with the background.
The final result:
In the generated video, the overall punching motion is fairly coherent, with relatively natural rhythm, without obvious "punch-retract" looping.
The background carries a chaotic street feel, creating visual contrast with the line-art figure. However, the detail layers of movement are insufficient — more a framework demonstration of martial forms.
2) Chinese-Style Anime Guy in an AJ Ad
One of the most popular play-around styles in AI-generated video is "Chinese-style anime."
So I had PixVerse V5 generate an AJ ad featuring a Chinese-style anime hunk, with simple plot and voiceover narration — a good opportunity to test PixVerse V5's "auto voice" feature too.
The prompt:
Nike AJ spokesperson video. A giant full-screen phone stands in an urban streetscape. A 25-year-old handsome man, 190cm tall, long legs, short hair, wearing white T-shirt, black tech jacket, dark gray straight shorts, AJ shoes. Standing before a giant full-screen phone, low-angle composition, giant phone standing in urban streetscape. Copy (deep voiceover): When trends no longer stay confined to screens — break boundaries, make the world step aside for you.
The generated video has an overall cool-tone atmosphere, visually leaning toward Hong Kong comic and CG styles. The voiceover timbre is fairly deep, with rhythm and pauses that connect reasonably well to the visuals, without obvious lag.
The result:
Due to conflicts between uploaded images and prompts, PixVerse V5 combined the creative elements. The overall effect is fairly complete, though some details remain noticeable — for instance, the character's expression is somewhat stiff.
However, from the integration of Chinese-style anime with the background, the coordination of visuals and voiceover already provides a usable template. The auto voice, while not strongly performative, is acceptable as an auxiliary tool.
3) Sci-Fi Aerial Combat in Realistic Style
After so many character anime attempts, let's try other styles.
Next, we see whether V5 can maintain compositional integrity and realistic detail in grand-scale imagery, fire and explosions, multiple subjects coexisting — especially in "realistic style."
The prompt:
A person shoots down a helicopter. Aerial overhead shot, sci-fi blockbuster quality, realism, panoramic, high altitude, warfare. Interior view from armed helicopter, soldier strafing ground with mounted heavy machine gun, shell casings flying, two armed helicopters firing small rockets at ground, multiple rockets with white trails flying from helicopter bays toward ground, ground covered in sea of flames with dense alien creatures, alien behemoths, explosions, black smoke, enhanced details, fill lighting, double exposure.
Such macro scenarios first test whether main subject actions and plot design, and special effects are sufficiently "realistic." Also whether non-main subjects become too templated, showing obvious "green screen" quality.
The final generated video shows armed helicopters, explosive fire seas, and [details]. You can see bullet spray, dynamic damage to helicopter bodies, with relatively clear local details.
Here's V5's generated video:
Overall, while the effects and details aren't flawless, within the AI auto-generation framework, the footage shows no obvious "green screen" quality or生硬拼接.
However, what's worth praising here is the matching between "sound effects" and muzzle flashes, gun mechanics — solid performance.
It's worth noting that the previous PixVerse V4 was essentially the industry's first "audio-visual integrated" AI video model.
4) First-to-Last Frame Splicing: Space of Text
This section's focus isn't on style but testing the "first-to-last frame" function — a foundational operation in many AI video generators. Creators can upload two comic-style static images and have the system generate a logically coherent short animation.

Users can first set frame count, i.e., how many images to upload. We'll start with the most basic "first-to-last frame" — 2 reference images.
For example, here are 2 calligraphy text-space images, with densely packed, repetitive, complex elements throughout the scene:

So I uploaded these two images to PixVerse V5, then input the story happening between them.
The prompt is actually very brief:
The protagonist walks out of the text barrier, only to enter an even larger space filled with text.
To match the visuals, I added a sound effect description:
Echoes reverberating through space
The entire generation took about 1 minute 30 seconds. Upon completion, the system produced four 4-second, 1080p resolution videos in one go. This generation consumed 360 credits.
I selected the best-performing video from among them to present here.
You can see that overall, visuals and audio combine quite naturally — for instance, when the protagonist crosses the text barrier and enters the next space, the transition is very smooth. And in the first scene, all the text maintains good three-dimensionality; as the character moves, there are essentially no hallucinations:
5) Multi-Frame Splicing: Dreamcore Glitch Electronic Music MV
Of course, "first-to-last frame" doesn't just turn comics into plotted animation — it can also accomplish stylized, stream-of-consciousness artistic creations.
My personal favorite style, and one heavily promoted across platforms right now, is "dreamcore." To explain: dreamcore isn't a single visual style but an aesthetic "between dream and reality," fairly abstract.
For instance, here I uploaded 3 dreamcore glitch-style images to see if they could generate an electronic music MV.

In multi-frame mode, you need to click in and set the "plot" and "timing" between each pair of images:

Since multi-frame setup is slightly more involved, I'll take a small shortcut and only input an overall prompt between the first 2 images:
An electronic music MV

The final result proves that PixVerse V5's comprehension of the "multi-frame" function is solid, with a complete final effect:
The generation didn't heavily modify the images, instead using camera push-ins and character movement transitions to connect them. For example, between the first image's figure and the second image's cells, a "gestation" process was established, giving the visuals an unexpected narrative quality.
It's admittedly somewhat abstract — but that's precisely the charm of "dreamcore."
Though the plot line remains thin, PixVerse V5 did achieve making seemingly unrelated images "look like a unified whole."
6) Continuation: Chinese Grotesque-Style Animation
Beyond "first-to-last frame," PixVerse V5 has another key feature: "continuation" — extending a video based on your uploaded footage and prompt.
The key here is consistency in character appearance and style across shots.
So I uploaded a Chinese grotesque-style video for evaluation. The original showed a puppet in Tang dynasty dress, with painted huadan opera makeup, holding an oil-paper umbrella. Being short, it had little plot.
Since it's Chinese fantasy grotesque style, we can try an interesting prompt:
The woman releases several duplicates of herself.

The continuation maintains the original video's Chinese grotesque style.
Character appearance and makeup remain fairly consistent; the extended duplicates, while differing in details, still show clear connection to the original figure. The atmosphere continues the original video's quality:
And, in case you didn't notice, the duplicate elements remain unified — like the pendant by the ear, which continues changing dynamically throughout.
The AI Video Industry Is "Rapidly Industrializing"
After testing so many cases, let's finally look at PixVerse V5's pricing and generation speed. In our tests, we mostly used 1080p, 8-second generations — 160 credits per generation, basically just over ten seconds, quite fast. The underlying technical architecture essentially determines this speed and its current largest user base.
From the PixVerse website, 15,000 credits = 459 yuan = 750 seconds; 1,000 yuan gets you 1,634 seconds. For annual membership subscriptions, domestic and international pricing has been unified with a 36% reduction.
This pricing, for someone like me who frequently visits various AI video generation platforms, is fairly "health-restoring."
Behind this lies the push of the AI video industry "rapidly industrializing."
AISphere was founded in April 2023; PixVerse opened beta in October. In the subsequent two years, it has had version updates nearly every quarter. From PixVerse V1 in October 2023 to V5 in August 2025, AISphere has updated its model version 8 times. In between, the domestic "Paiwo AI" only launched this June.
Two months live, and product features, modules, performance, and cost-effectiveness are fully aligned with the international version.
AISphere's technology has consistently been doing "subtraction."
For example, V3 in October 2024 launched no-prompt special effects templates, letting many people realize "so I can make AI videos too" — so more people played with it; by V4, video generation speed increased, the app launched, and users began treating it as an everyday tool; then V5, with substantially improved generation quality and efficiency, plus the Agent creative assistant, lets zero-experience users turn a single image into a narratively complete short film — truly upgrading from "just playing around" to "actual creation."
This gradually formed a "more users → more content → even more users" snowball effect.
Many creators joke that PixVerse's (Paiwo AI's) iteration speed resembles early mobile game studios. Not afraid of many versions — afraid of slow updates.
What's the concrete manifestation of cheap and fast updates?
Take the simplest example: if you scroll short videos daily, you've probably noticed AI short videos and AI short dramas rapidly emerging in major traffic pools. In our own recommendation feeds, one in several posts is AI-generated. The driving force behind this is very "industrial" — creators need "industrialized" AI video tools: cheap, easy to use, stable, fast output:

It could be said that when an AI video generation model can advance further across four dimensions — "real-time, story-aware, aesthetic, consumer-facing" — AI video generation "industrialization" will have had a proper beginning.
🚥
After this dizzying tour of tests, you might feel it's all rather complex. But if we strip away all the technical jargon, the core of this matter is actually very simple, and full of idealistic warmth:
It tells us that the barrier to video creation is disappearing at an unbelievable speed.
PixVerse's (Paiwo AI's) co-founders possess a unique "sense of the lens." They frequently participate in major conferences, using communication opportunities to expand their influence.
Founders Changhu Wang and Xuzhang Xie have both expressed their views on the future of AI video generation at various occasions:
Let good models bring good products.
There are still billions of people worldwide who've never made a video. We hope to use AI to help these vast majorities achieve universal access to video creation.
These statements mean that where we were once audiences to stories, going forward, everyone has the opportunity to become a storyteller.
Every unique, sparkling idea in each person's mind can be captured by the "AI camera."

