No Hype, No Hate: Kunlun Tech's New Model Is Kind of a Big Deal

The business world has never gone easy on "good technology."

The business world has never been kind to "good technology."

👦🏻 Author: Jingshan

🥷 Editor: Koji

🧑‍🎨 Layout: NCon

If we wanted to, we could easily rattle off a long list of "big names" — every person, every technical approach angling to become the "top player."

The business world has never been kind to "good technology."

From virtual streamers in livestream rooms to digital avatars fronting brand campaigns, the market is waiting with unprecedented hunger for these technologies to break through.

Starting in 2024, we distinctly felt domestic players beginning to "flood in" to this space. There's Tencent's open-source HunyuanVideo-Avatar; and Kunlun Tech, iterating rapidly from SkyReels V1 through SkyReels V2, each offering their own answer.

Just yesterday, Kunlun Tech's SkyReels-A3 model officially launched, with a slogan that's simple and punchy: "Audio-driven digital human video."

"Audio-driven digital human video"

🚥

As the Crossing team, we weren't about to miss this chance to be early testers.

We ran 10 cases, from real-person interviews to binary-code portraits, even getting a chibi Mona Lisa to speak. We also dug through the relevant technical report, finding 5 key highlights that explain these performance results.

Below, the complete test records and technical breakdown.

10 Cases: From "Real People" to "Binary-Code Portraits"

Simply put, SkyReels-A3 is an audio-driven video generation model. It only needs one image and one audio clip — no prompt engineering required — to produce solid digital human video results.

It performs well across interviews, singing, short-form marketing, podcasts, and many other domains.

Person is talking

Since this model is genuinely fun to play with, the Crossing team tested 10 cases total. We spent a while debating what order to present them in. In the end, I decided to sort them by increasing "abstraction" of the uploaded image.

Though somewhat counterintuitive, realistic human portraits are actually easier for SkyReels to handle.

Let's go through them one by one.

1) Sam Altman's Classic Interview

SkyReels-A3 really does need just one image plus one audio clip to generate remarkably natural, fluid video. And in this 40-second clip, the figure maintains very natural movements synchronized with the speech output.

For example, I uploaded a classic interview photo of Sam Altman — one that many people have probably seen countless times by now.

In this photo, Sam Altman tilts his head toward the interviewer. Note the hand position:

Then I grabbed a random audio clip from some program.

After combining and uploading, SkyReels-A3 generated the video below. If you look closely, you'll notice that in the very first second, Sam Altman has a pronounced breathing motion (his chest rises and falls).

When emphasis hits (like on "Truly"), his hands even reach forward in response...

Brief pauses like "Umm" are precisely rendered in Sam Altman's mouth movements, and even his exasperated smile shows completely natural lip shapes.

2) The Female Streamer Selling "Big Orange" Pillows

After that Sam Altman test, I felt SkyReels-A3 is incredibly well-suited for livestreaming and short-form marketing. So we tested how a female streamer performs under "livestream fill light."

Given copyright concerns, our subsequent materials were mostly generated from various AI tools.

Below is a typical Douyin livestream scene: a young, attractive female streamer in a streetwear hoodie, smiling at the camera while presenting a product — a large orange pillow:

Let's look at the video generation results.

First, the lip sync is very natural. Mouth movements match the speech with high accuracy, with smooth lip shape changes during speech, basically no noticeable delay or misalignment — the viewing experience feels quite real and fluid.

Actually, there's a very effective standard ordinary people use to judge whether a digital human video looks too fake: whether there are hand movements, how varied the hand movement patterns are, and whether they look natural.

In this 39-second video, you can count how many hand gestures the streamer makes:

Finally, one detail I find quite impressive: the streamer's left hand is often behind the orange pillow (green area), and when she extends it, there's almost no deformation in the image.

3) Letting Koji Introduce His Background

Next, we invited Koji to talk about his personal history. Honestly, this case felt the most realistic to us.

And because we all know each other well, seeing the video felt even more real at first glance.

I grabbed Koji's personal photo from a random video on his WeChat Channels, and the audio too:

This generation had one fairly noticeable minor flaw: the "Crossing" text on Koji's shirt deformed. Of course, this is also a common weakness of current AI models in Chinese text generalization.

But when it comes to the naturalness of Koji's language and facial expressions in the video, and the variation in gestures — friends who know him can probably appreciate just how realistic this is.

Especially in some emphasis-heavy moments, the digital Koji differs somewhat from Sam Altman: this time, the emphasis comes mainly through intensified facial expressions:

After seeing how real-person photos perform in SkyReels-A3, let's look at "the stuff we messed around with." During testing, I found that SkyReels-A3 is especially suited to diverse characters and multiple subjects. Its biggest strengths: strong generalization capability, and stability.

4) Modern Urban Anime Guy

Below is a modern anime-style image. The facial features are completely different from real human facial details, and the light patterns behind him might make it even harder for an AI model to recognize.

Here's the video generation result. We can still see quite good lip-to-voice synchronization.

Not sure if you noticed a detail: the anime guy's teeth match the overall style of the image, and his chest also rises and falls with the audio.

The light patterns on the character's body also shift, his hair even has a sense of "weight," and his eyes continuously blink:

5) Chibi Mona Lisa

The next character might be even harder for AI to "recognize."

Like this 3D moe-fied Mona Lisa below, in a cute hybrid style somewhere between Pixar and Pop Mart. The face is round and adorable, chubby, with eyes in that distinctly anime-style gradient star pattern. Note on the left: little Mona Lisa also extends a chibi hand.

Now, let's have her introduce the "Crossing Podcast."

SkyReels-A3's results remain quite good. Pupil dilation and constriction show realistic reactions, and the picture frame behind her doesn't deform.

On the line "Welcome to our podcast," she even extends her little hand. Overall style consistency is quite good, maintaining the Pop Mart and Pixar aesthetic.

Little Mona Lisa's eye blinking is genuinely cute — it doesn't default to "human blinking" just because it's a chibi character:

6) Middle-Aged Man in Charcoal Sketch Technique

The image below ratchets up the difficulty again.

This is a portrait rendered in charcoal sketch technique, in a "neo-Chinese" comic style with traditional ink-wash sensibility. The color palette is basically just black and white:

Let's have him try singing my personal favorite, Freddie Mercury's "Bohemian Rhapsody," and see how it goes.

The opening line's presentation is already somewhat startling:

Throughout the video, if you look closely, the character's facial expressions, neck tension, even his Adam's apple all shift with the emphasized words in the song.

And throughout, the figure doesn't stop natural movement despite being a hard-to-recognize "charcoal sketch."

7) Conan Circular Badges

Next, let's see what happens when the image contains more than just a single human figure.

Like this Conan circular badge composition, made entirely of black, white, and gray tones, with many instances and complex elements:

Here's the final generated result.

For most of the video, only the centermost Conan badge speaks, and while outputting audio, the badge also rotates.

In the latter part of the video, the two side badges also start "blinking," creating an overall effect with some "uncanny valley" vibes:

The cases above at least had some recognizable "human figure." But the following sections "can barely be said to have a recognizable facial structure."

8) Letting Gundam Introduce Itself

Since SkyReels-A3 could even handle Conan badges, I decided to let it try "Gundam" and see what would happen.

Below is a frontal close-up of a Gundam head, with clear head structure. Note the small patch of metal armor surface below the head, which also has some gloss and reflective variation:

Here's the video result.

First, a minor flaw: in the first 10 seconds, SkyReels-A3 actually auto-generated a pair of "mechanical eyes" based on the Gundam image — which I thought looked pretty good. But after a few seconds, they automatically disappeared.

Looking more carefully at the video details, I noticed the reflective changes on the metal armor below the head are quite realistic and natural, and the whole head naturally sways while speaking.

9) Modern Abstract Art Face

Finally, two "abstract paintings."

Honestly, even for humans to identify, it's hard to imagine "what would happen if this became a digital human."

SkyReels-A3 gave its own answer.

Below is a futuristic modern abstract art poster, blending fashion, technology, and stream-of-consciousness aesthetics. The image consists of countless colored particles and code symbols, with a vague human face outline:

Let's have this "abstract painting portrait" sing GAI's hip-hop classic. Video results below.

Most noticeable is the dynamic particle performance behind the portrait, with rich visual layers. Additionally, the stream-of-consciousness face changes fairly naturally with the audio transformations, especially in the facial area around the lips.

10) Binary-Code Portrait

One more even more abstract image: a binary-code composition showing an abstract profile silhouette. Note in the lower left, this binary-code figure also extends a hand:

Let's have it introduce DeepSeek's short-form marketing script. Final video results below.

You can see in the lower left corner, the binary-code figure's hand also continues to change with the audio. Honestly, I didn't expect this binary-code portrait's mouth to look like this as a digital human video, but it is indeed fairly natural:

During the experience, you can clearly feel that the SkyReels-A3 model is genuinely audio-driven.

Even without writing complex prompts, simply telling the AI "this person is talking," the entire model automatically makes the digital human's mouth, expressions, even gestures change with the audio content.

This "audio-driven fluidity" also sparked strong curiosity about the technology behind this model, so I went searching on arXiv for the SkyReels-Audio technical paper — and indeed found some rather unusual technical innovations by the team behind it.

Next, let's dive deep into these highlights.

We Found 5 Highlights in SkyReels-Audio

SkyReels-A3 is essentially a DiT-based (Diffusion Transformer) video generation model, which is currently the mainstream model framework, specifically designed for the "generate video from sound" scenario.

Its biggest characteristic: it can generate temporally coherent, visually consistent long-form video content based on audio, while also precisely controlling the intensity of expression in each frame.

For example, when the audio calls for emphasis on a certain phrase, the digital human's expression intensifies accordingly; where speech is fast, the movement rhythm keeps pace. This control capability is precisely the result of the technical team's breakthroughs.

To achieve this technical goal, the team made a series of breakthroughs.

We'll mainly discuss how these technical breakthroughs "cash out" into the model's actual performance.

1) "Can Keep Talking"

Most current AI models basically focus on generating 3-5 second short videos. If you try to go slightly longer, things are likely to "crash" midway: either the frame suddenly freezes, or there's clear discontinuity with obvious stitching痕迹.

SkyReels-A3 has partially solved this "endurance" problem.

It uses a training-based video extension method. By training the model to simultaneously accept both a "reference image" and already-generated "historical information segments," it ensures video continuity while maintaining image quality and character consistency.

Sounds technical, but think of it this way: it alleviates common problems in long-video stitching (tearing and quality degradation), enabling near-"infinite length" coherent output.

This SkyReels-A3 can "connect" speaking videos very smoothly, maintaining coherence over relatively long durations without breaking down.

From the examples provided by the technical team, videos generated using the training-based conditional fusion method show significantly reduced ghosting and better motion continuity.

2) "Full Modality"

The current SkyReels-A3 establishes a "full modality" audio-driven framework, meaning: it doesn't just understand sound, but can simultaneously comprehend text descriptions, reference images, and video content.

All this information is processed within the same technical framework, capable of both generating entirely new speaking videos and editing existing video content.

This mainly owes to the team's unified integration of audio + text + image + video for conditional control on the same Video Diffusion Transformer, targeting both generation and editing tasks for "speaking human portraits."

3) "More Accurate Lip Sync"

For digital human video, lip synchronization is the most critical technical challenge. Human eyes are extremely sensitive to "mouth not matching words" — even slightly off feels strange.

Not every part of the face contributes equally to the "speaking effect."

SkyReels-A3 invested heavily here.

Focused Attention on Mouth Region

The team introduced a "facial region weighted loss" mechanism.

Simply put: during training, the model is made to pay special attention to lip region changes, while maintaining normal attention to relatively static areas like the forehead and nose. This achieves precise lip detail without making the entire face look rigid.

During video editing, the team employs a two-stage strategy: early steps use the full video to maintain structural consistency, while later steps switch to the first frame image to reinforce lip detail.

More Natural Fit

It integrates Whisper audio feature extraction technology, using special fusion methods to make lip shapes fit sound more closely, for more natural speech. Simultaneously, during inference it adds Audio-Guided CFG (audio-guided sampling weights), dynamically balancing conditional influence to strengthen "audio-visual synchronization."

In the technical report, SkyReels-A3 also provides subjective effect comparisons with other audio-driven talking head methods, showing clear advantages:

4) "Beyond Accuracy, It Needs to Look Like Them"

SkyReels-A3 employs a "hybrid learning" training approach, meaning: two training methods learned together, with complementary strengths.

For example, the model uses two types of training materials simultaneously:

[1] One type: one photo + one voice clip → generate talking animation;

[2] The other type: one reference video + new voice → edit lip movements to make it say new content.

These two methods each have strengths: photo-driven animation emphasizes "speaking well," with more prominent lip detail; while video-driven editing emphasizes "looking like the person," making facial expressions more coherent.

During actual training, SkyReels-A3 trains them jointly.

This dual training makes the final results both accurate and natural.

5) "The Dataset Is Good, and Strict"

Everything above is basically architectural improvements. Beyond that, SkyReels-A3 also needs quality training data.

Looking at their data flow diagram, the screening does feel genuinely "solid."

SkyReels-A3's training data undergoes strict selection. As the paper states:

From 10,000 hours of raw video, 1,000 hours of high-quality training set were carefully selected. This process includes multi-stage quality checks and human annotation, ensuring every piece of data meets standards.

SkyReels-A3 Data Processing Pipeline

Additionally, the team built their own test benchmark covering "50+ scenarios, multiple languages, multiple speaking styles" to ensure robust performance.

From the team's provided test results, SkyReels-A3 does perform quite well in audio-lip alignment. We also pulled a few comparison images of SkyReels-A3 against other baseline methods in audio-lip alignment:

Interestingly, as we tested earlier, it can also handle various different character styles, including cartoon characters.

The team also tested with reference images of different characters, sizes, and styles, and the final video results were relatively natural and consistent.

For example, Peppa Pig's speaking video shows very natural lip sync, well-matched to the original character's features.

After 10 cases and one technical report breakdown, we've deeply felt one fact:

Technology is indeed galloping forward at speeds beyond our imagination.


Looking back, just one short year ago, we were still marveling that AI could generate a few seconds of video without "breaking character." Back then, a "Will Smith eating noodles" video could draw collective gawking and heated discussion across the entire AI video field.

And now? Now, AI can already have digital humans tell complete stories, express rich emotions, with lip sync nearly perfectly matched to voice.

The Crossing team will continue tracking progress in AI video digital humans, mapping out this "static to dynamic"赛道.