HappyHorse 1.0, your happy little pony is now online
Happy Pony, on Qwen.
HappyHorse, Now on Qwen

👦🏻 Author: Jingshan
🥷 Editor: Koji
🧑🎨 Design: NCon

Alibaba's ATH (Alibaba Tongyi Lab) has officially announced that HappyHorse 1.0 — "Happy Little Horse" — will begin closed beta testing on Qwen starting today.
Two weeks ago, an AI video model with no official introduction, no launch event, and even a somewhat casually chosen name appeared on the Artificial Analysis leaderboard: HappyHorse 1.0, or "Happy Little Horse," scored an Elo of 1349 in the text-to-video without audio category, and 1403 in image-to-video — stunning everyone.

The internet immediately started guessing who made it. The discussion spread to Reddit and X, with everyone poring over its generated samples, watching and speculating, until Alibaba ATH stepped forward to claim it.
Today, HappyHorse 1.0 has officially entered closed beta.
🚥
The Crossing team secured beta access and immediately put HappyHorse 1.0 through its paces, sharing our hands-on experience and observations below.
After extensive testing on Qwen, here's our conclusion first: the word that best describes HappyHorse 1.0 is "stable."
Across several different scenarios we designed, with multiple runs each, the results remained consistently reliable.
Some basic specs: it supports both text-to-video and image-to-video, with maximum resolution of 1080p, duration options of 5, 10, or 15 seconds, aspect ratios of 16:9, 9:16, and 1:1, and optional audio.
At its core is a 15-billion-parameter unified Transformer model that generates audio and visuals together — architecturally distinct from approaches that produce silent video first and dub it separately. It also supports multilingual lip-sync, covering Mandarin, Cantonese, English, Japanese, Korean, and more.
Qwen has two beta entry points: the Qwen app and the Qwen Creation web interface (c.qianwen.com). If you've been selected for the beta, you can start using it right away.
Below, we walk through five specific scenarios and what we observed.
Five Tests: What We Saw
Case 1 | Camera Language
As an audio-visual integrated model, HappyHorse 1.0 can output videos with sound directly, supporting multiple languages including Chinese, English, and Japanese. It performs steadily in dynamic, moving-camera shots, maintaining basic character consistency throughout push-in movements.
Prompt:
**"Early morning on Shanghai's Bund, thin mist still lingering. The camera pushes horizontally from the river surface, passing elderly Chinese and American people doing morning exercises, finally settling on a cup of steaming coffee by the roadside. Strong sense of depth throughout, light transitioning from soft to backlit. 1080p, 16:9, with ambient sound."**
Looking closely, the elderly person in pink at the front keeps moving while speaking, with fairly natural facial animation. The man on the right pauses to put his hands on his hips before turning to speak to the person on his left. Between camera transitions, characters adjust their movements accordingly, with overall visual consistency maintained.
The river view behind the Bund and the iconic Pudong skyline beyond match real-world spatial relationships.
Case 2 | Livestream E-Commerce
This next scenario is more abstract, and also quite representative: livestream product selling.
Prompt:
**Fixed camera. An Asian female host in a floral dress sits at a white desk in a livestream room decorated with pink-purple neon lights, holding a pink skincare essence bottle. She smiles and looks directly into the camera introducing the product. Background features pink curtains and neon rocket decorations. Multiple skincare products and a makeup mirror are arranged on the desk. Overall aesthetic follows beauty blogger livestream style, 4K HD quality, soft and bright lighting. Spoken content: "Babes, this miracle essence is absolutely amazing! Hydrating and brightening, leaves your skin glowing and dewy. Special price today in the livestream, go grab it now before it's gone!"**
This scenario demands more from visual performance. The host needs to introduce products while corresponding gestures and lip movements change, but these are all rendered fairly well with decent consistency. Audio-visual sync generally holds up.
Looking closely, you'll notice all cosmetic product logos consistently display "Qwen" in English text. The lighting effects on the cosmetics feel authentically realistic.
Case 3 | David Fincher Style
Hollywood commercial film color grading and texture demands are quite specific. This test focused on precision in executing color directives and sound design approaches.
Prompt:
"Dark green dominant tone, high-contrast hard shadows, underground parking garage, flickering cold white fluorescent lights, handheld 35mm cinematography, medium close-up follow shot. A man moves quickly between vehicles, footsteps rapid and slightly stumbling. Sound design: heavy breathing close to microphone, strongly reverberant footsteps, man's low muttering intermittent and suppressed. 1080p, 16:9, with audio"
The color direction is correct — shadows lean green-blue, highlights lean cold white, overall contrast is strong. The fluorescent flicker comes through; it's not steady constant light, with occasional single-frame flashes. The parking garage texture feels right.
Handheld shake is present — slight vertical jostle as the character runs, conveying follow-shot tension. The sound design layers three elements: breathing in the foreground, reverberant footsteps with spatial depth in the mid-ground, and occasional muttered words buried underneath. The three layers remain fairly distinct.
Looking closely at the light interaction where the man's hair overlaps with the overhead lighting, the physics check out:

Case 4 | Hong Kong Style
HappyHorse 1.0 actually suits Hong Kong-style short films quite well. Whether it's old-movie texture or more realistic TV drama visuals, it handles both.
On one hand, the visual style is steady. On the other, its understanding of the real world is decent — it doesn't easily go off the rails. So I directly tasked it with creating a classic TVB crime-comedy short, referencing Forensic Heroes.
Prompt:
`Classic TVB crime drama visual style, reference Forensic Heroes and Armed Reaction (serious cold lighting, B holding notebook, A speaking in lowered voice)
A: (serious) Got a tip. Big Ming raised Little Ming. Tell me, who raised Big Ming?
B: (taking notes seriously) His father?
A: (cold shake of head) Wrong.
B: (nervous) Could it be... the murderer?
A: (word by word) It was a dog.
B: (shocked) Huh? Why?
A: (suppressed laughter breaking through) "Gou yang" (dog raised) Big Ming, get it?`
Overall, the workflow went smoothly this time, with fairly realistic final results.
Looking at certain clips, the dialogue, tone, and especially speech rhythm match the classic TVB crime drama style specified in the prompt. The characters' overall posture and clothing fit hit series like Forensic Heroes.
This indicates it draws on some world knowledge. It knows roughly how this style should speak, perform, and shape overall narrative direction.
Case 5 | Tang Dynasty Group Portrait
The final test was a multi-person group scene — a persistent challenge for AI video, with many characters, many movements, complex lighting, and frequent issues like merged faces or animation glitches.
Prompt:
**"Tang Dynasty-style inn, wooden structure, carved window lattices, red lanterns hanging, mixed candle and oil lamp lighting, slight haze in the air, layered light distribution. Over ten people seated at several tables, wearing Tang-style robes with wide sleeves and cinched waists. Some slightly tipsy, some laughing loudly, some expressionless. They're playing finger-guessing drinking games. 1080p, 16:9, with audio"**
In the multi-person scene, each character's state is differentiated. The camera gradually pushes in, with smooth transitions between multiple shots and overall logical coherence. HappyHorse 1.0 works well with short prompts, letting it fill in scene development on its own.
Costume and prop details also deserve mention. With over ten people seated at several tables, each person's robe patterns, waist cinch positions, and wide sleeve cuts vary — no sign of the entire table sharing identical texture maps.
Looking closely at some characters, HappyHorse 1.0 maintains fairly stable performance in human figure realism.
Here's a still frame — setting aside movement, the character rendering alone shows strong fidelity:

The AI Video Leaderboard's Musical Chairs
The pace of AI video model development is worth pausing to consider.
A timeline makes the rhythm clearer.
May 20, 2025: Google unveiled Veo 3 at I/O, becoming the first mainstream video model to integrate audio-visual generation. Late that same month, Kuaishou's Kling 2.1 launched, emphasizing 1080p HD and cost-effectiveness; by its tenth month, Kling AI had already surpassed $100 million in annualized revenue. September 30: OpenAI released Sora 2, with its iOS app hitting one million downloads in five days, trending across overseas social platforms and pushing the Sora App to the top of the US App Store overall charts.
October 15: Google quickly followed with Veo 3.1. Entering early 2026, Kling successively released 2.5 Turbo, O1, and 2.6, making audio-visual synchronized generation a baseline capability.
Then on February 12, 2026, ByteDance's Seedance 2.0 topped Artificial Analysis with an Elo of 1,269, and HappyHorse 1.0 appeared on the same leaderboard shortly after. From half-year rotations, to months, to weeks this time — the cadence is compressing.
Each leadership change, the interval between them grows shorter.
But one thing needs clarification: these "number one" swaps involve score differences that aren't actually that dramatic. Top-tier models each have their strengths across different scenarios; some hold price advantages. It looks like generational iteration, but a more accurate description might be:
Everyone's capabilities cluster around a similar level — musical chairs, frequent swaps, but the gap is narrowing. Even temporarily "surpassed" predecessors are still expected to potentially release a new SOTA model and reclaim ground at any moment.
So for users, which to choose today depends largely on what exactly you want to do, and how much time and money you're willing to spend. For content creators, this rhythm feels somewhat complex. Just when you've gotten comfortable with your current tool, something better may appear next month; chasing every update is exhausting, but not chasing risks missing out.
Fortunately, one thing is happening in parallel: the barrier to accessing these models is genuinely dropping — as with HappyHorse 1.0's simultaneous multi-platform beta on Qwen this time.
Strong model capabilities plus nearby access points — these two factors combined create real impact. In AI video competition, impressive demos are only part of the equation.
For content creators, the most practical change right now is: pick up your phone and try it, no waiting, no payment. See if it's useful, then go deeper. AI-generated video has in fact reached the stage where "anyone can give it a shot."
How close the entry point is determines how fast this process unfolds.
🚥
The story of "Happy Little Horse" galloping forward has only reached its halfway point.

Crossing is seeking independent contributors to write AI product and model reviews.
If you've written articles like "Hands-on: PixVerse C1" or "Hands-on: LibTV", please contact zeo0811@gmail.com. Your email should include: ① personal introduction, ② AI review articles you've written.
We offer competitive compensation. We look forward to observing and documenting the AI era together 🎪
