When Two "World Number Ones" Appear at the Same Time — Written After Hunyuan and Keling AI Both Claimed the Global Top Spot

What a time to be chasing each other!

What a time to be alive — neck and neck!

👦🏻 Author: Jingshan

🥷 Editor: Koji

🧑‍🎨 Layout: NCon

Recently, two announcements dominated tech headlines: on September 23, Kuaishou revealed that its Keling AI 2.5 Turbo image-to-video and text-to-video model, just ten days after launch, had claimed the #1 spot globally on Artificial Analysis; hot on its heels, Tencent announced that its Hunyuan Image 3.0 model had reached #1 on LMArena.

Hunyuan Image 3.0

Kuaishou Keling AI 2.5 Turbo

What's fascinating is that Kuaishou's Keling AI 2.5 Turbo and Hunyuan Image 3.0 each topped different leaderboards, yet both were conspicuously absent from each other's rankings. Kuaishou's name didn't appear on LMArena; Tencent's name was nowhere on Artificial Analysis.

So what do these two "world #1" titles, born almost simultaneously, actually mean? Does a model reaching the top of a specific leaderboard signal comprehensive superiority over rivals — or simply that it's better at gaming a particular set of rules?


Next, we'll reconstruct the full picture of this "leaderboard war," break down the rules and logic behind these rankings, and understand what "being #1" actually signifies in today's AI landscape.

How exactly is "first place" calculated?

To grasp the meaning of "who's #1," we need to unpack the rules of the "leaderboards" or "arenas" themselves.

AI model rankings are products of specific evaluation systems. "First place" depends on which arena a model won in, and under what rules.

Let's examine the two representative leaderboards where Hunyuan Image 3.0 and Keling AI 2.5 Turbo claimed their crowns: LMArena and Artificial Analysis.

LMArena

First, LMArena's text-to-image leaderboard (the one Hunyuan 3.0 used).

LMArena is essentially a human preference voting "arena" launched by UC Berkeley, and currently one of the most authoritative international leaderboards.

Put simply, it doesn't care about parameter counts, inference speed, or technical metrics — it hands the final judgment to thousands of real human users.

Its rules are, overall, "brutally simple," like an anonymous "pick one of two" blind test.

The rough flow:

【1】A user enters a prompt;

【2】The platform randomly sends this prompt to 2 anonymous models, each generating an image;

【3】Both images appear side by side, and the crowd votes for which one they prefer.

A simple example:

Model A (GPT-5) vs. Model B (Claude 4.5): after viewing both responses, the user votes: B is better. If GPT-5's current Elo is much higher than Claude's, B's win counts as "defeating a stronger opponent," so its score rises significantly. If B keeps beating high-Elo opponents, its Elo keeps climbing. When it repeatedly defeats strong rivals and performs consistently, its Elo gradually converges toward its true skill level.

When a model defeats a higher-ranked opponent, its score jumps more; conversely, less. After massive volumes of user votes, the system calculates a score for each model representing its relative position in popular aesthetic preference.

To be more rigorous, LMArena's scoring system involves extremely complex "pairing strategies, qualification mechanisms, style control," and more.

For instance, to focus votes on "how well models understand prompts, content, and expression quality" rather than flashy styling, LMArena researched how to dampen interference from "format, output style, layout, Markdown, and other non-content preferences," introducing modeling and calibration mechanisms for stylistic and formatting factors.

Another example: to make score updates more efficient while preventing some models from going un-compared for too long, they adjust selection probability based on current score uncertainty (confidence interval), ranking differences, and other factors.

It sounds incredibly complex.

So, strictly speaking, this demands a serious academic paper to support their methodology — and LMArena has published exactly that:

Therefore, Tencent's Hunyuan Image 3.0 reaching #1 on LMArena's global leaderboard means that among 26 top-tier models in close combat worldwide, Hunyuan earned the highest score in the test of "which image real users found more appealing."

Artificial Analysis

Next, the Artificial Analysis video generation leaderboard where Keling AI 2.5 Turbo won.

Similar to LMArena, Artificial Analysis also employs sophisticated mechanisms, blending traditional benchmarks with "comparative voting"-style systems.

Its mechanisms broadly divide into two directions based on what AI capabilities are being tested:

【1】For text, language, understanding, and reasoning, it has an "AI Index" (Artificial Analysis Intelligence Index, AAII), composed of multiple sub-benchmark tasks.

These include reasoning, knowledge Q&A, mathematics, coding, long-term memory, and contextual understanding.

【2】For image and video generation, it has dedicated "comparison mechanisms" for blind-test-style model comparisons.

For example, its Video Generation Arena allows users to compare videos generated by 2 models from the same prompt.

Such as:

【1】Motion smoothness and physical realism:

Does the video move fluidly and coherently, following physical laws?

【2】Object and identity consistency:

Do people and objects in the video suddenly change appearance mid-stream?

【3】Aesthetic quality:

How's the composition, lighting, and overall visual impression?

【4】Semantic alignment:

Does the video content match the prompt? (After all, prompt text is provided during blind comparisons.)

And so on.

It was on this technical track that Keling AI 2.5 Turbo 1080p defeated trailing SOTA models on both text-to-video and image-to-video sub-leaderboards, taking the crown.

While browsing LMArena and Artificial Analysis, we noticed a thought-provoking phenomenon.

Two models performing at the top of their respective leaderboards — Hunyuan 3.0 and Keling AI 2.5 Turbo — were both "absent" from each other's rankings. Specifically, Hunyuan 3.0, which topped LMArena, did not appear on Artificial Analysis's image generation leaderboard; and Keling AI 2.5 Turbo, which claimed "double crowns" on Artificial Analysis, was likewise not included in LMArena's text-to-video or image-to-video leaderboards.

If you have insight into why, please share in the comments!


Beyond them, more "arenas" are running

In reality, LMArena and Artificial Analysis are just two representatives in AI evaluation. To more comprehensively measure AI capabilities, there's now a much larger "tournament system."

The evolution of this evaluation system itself is a microcosm of innovation in the AI field. As competitors grow stronger, evaluation systems must keep pace.

Initially, everyone relied on fixed academic benchmarks; later, as generative AI answers grew more open-ended, platforms like LMArena that incorporate real human preferences emerged; then, to prevent models from learning "test-taking tricks," anti-cheating "closed-book exams" like LiveBench were introduced.

Now, technical evaluation alone lags far behind the times, and "humanistic arenas" like HumaniBench focusing on ethics and fairness have begun appearing.

This actually shows that our definition of "good AI" is expanding from single-dimension technical quality to a more multidimensional framework.

So, if leaderboards are deliberately designed exams, then an AI model reaching the top doesn't always mean it's "comprehensively stronger" in all aspects — it may simply mean it's better at taking that particular "test."

Much academic research has explored this. Nature published an article with a hard-hitting title: Is Your AI Benchmark Lying to You? It pointed out that over-pursuing high scores on specific benchmarks is distorting our understanding of AI's true capabilities.

There's abundant research like this we won't enumerate.

Returning to more concrete scenarios, as the most popular "arena," LMArena has faced intense debate around fairness. Critics have long suspected that major companies' models undergo extensive "gray testing" on the Arena under codenames before official release — giving them several extra rounds of "test prep."

The-decoder published an analysis based on over 2.8 million model comparison records (January 2024 to April 2025), suggesting large model vendors may leverage greater resources, more frequent version submissions, and debugging opportunities to gain asymmetric exposure and improvement on the platform.

But since LMArena consistently denied this, no one caught any "hard evidence," leaving people skeptical.

Then Meta's Llama 4 walked straight into the controversy, making everyone fully believe that "major companies do optimize for leaderboards."

Here's what happened: In April, Meta uploaded a version of Llama 4 called "Maverick" to LMArena, immediately scoring second place.

However, observers quickly noted: the model version submitted to LMArena didn't fully match the version publicly released to the developer community. Reports indicated Meta used an "experimental chat version" for evaluation — one optimized for conversation.

TechCrunch and The Verge directly reported:

If a model is specifically tuned to perform better on this benchmark, but the version ordinary users receive lacks that optimization, the benchmark score is misleading.

Users noticed the Llama 4 LMArena version seemed to use excessive emojis and gave very verbose responses.

Many viewed this as a textbook case of "score-showcasing via optimized version."

Of course, LMArena's official team continuously evolves "exam rules," working to reduce interference from these non-capability factors.

None of this discussion aims to negate leaderboards. On the contrary, "exams" are necessary catalysts for progress — they let us clearly see AI models' improvement under specific rules, specific datasets, and specific user populations.

The real competition isn't about being #1

In the evolution of AI models, no "first place" has ever stayed at the top forever. From the "ancient" era of DALL·E to Midjourney, from GPT-3 to GPT-5, to Sora, Hunyuan, and Keling.

Every era has its champion, and every champion is destined to be surpassed.

The core value of this race has rarely been about capturing a fleeting rank. Its core value lies in the process of neck-and-neck competition itself, making all participants faster, better, and more open.

The true arena lies beyond the leaderboards. It concerns four more fundamental dimensions — we've organized these briefly, and welcome your additions:

【1】Generalization capability

What we often call "learning by analogy" — AI models can't simply memorize "standard answers."

There's a fascinating case that illustrates this perfectly.

NeurIPS 2023 featured a dramatic fine-tuning competition where the first-place model in the public phase saw dramatic performance drops in the closed-book phase. This is the classic example of "good at tests, not necessarily good at learning by analogy."

【2】Robustness (what we often call stability)

Facing messy real-world inputs, will the model easily "break" or produce harmful outputs? It needs to gracefully handle all kinds of "uncertainty."

【3】Cost efficiency

An evergreen issue: intelligence comes at a price, and that price is massive compute. How to maintain high performance while reducing training and inference costs is key to whether technology can move from lab to market.

Nearly all vendors are fiercely "competing" here. One striking example: when OpenAI released GPT-5 in August, it immediately pulled GPT-4o and other models entirely. Because it was smarter, while costs stayed flat or dropped.

The impact on user and market experience was enormous.

【4】Multimodal fusion

We once made this assertion: the more modalities an AI model fuses, and the higher the quality, the better. Because future AI will be a seamless综合体 of listening, speaking, reading, seeing, and thinking — an "intermediate layer," more like an OS.

It needs deep understanding of connections between different modal information.

Like this time's champion Hunyuan 3.0 — much attention focused on its existing "world knowledge" and how it applies this information in image modality generation.

Hunyuan 3.0 generation

Neck and neck — what a time to be alive

If we were to write one sentence summarizing the AI competitive landscape from 2025 to now, it would undoubtedly be that phrase the Crossing team has returned to again and again:

"What a time to be alive — neck and neck."

We cannot, and need not, pursue an eternally unchanging "first place."

Because the essence of AI is dynamic evolution. An AI model that seems perfect today may be surpassed tomorrow by new architectures and new algorithms — just as Transformer and DiT architectures emerged. This is precisely what makes this field most thrilling.

From initially generating blurry images to now creating videos indistinguishable from reality. Behind this lies countless teams' day-and-night R&D, open community contributions, and the potential unlocked by competition.

Therefore, we're proud not only of the Hunyuan and Keling teams, but we cheer for all researchers, engineers, and creators still on the journey.

Perhaps this is what "first place" truly means in the AI era:

It's not about defeating others, but about proving to the world that we have one more better, stronger, more ideal answer.

And as long as this virtuous chase continues, the best answer is always the next one.