To Video Models: If the Masses Don't Love You, Who Do You Think You Are?

**"🤡meme bench Released"** To better evaluate video models, ZangAI spent half a month on a big project: meme bench.

"🤡meme bench Released"

To better evaluate video model capabilities, Zang AI spent half a month on something big: meme bench.

Most people don't realize a blind spot here — whether a large model is "good" is a matter of definition. Benchmark is power.

ByteDance's video model Seedance 2.0 was undefeated for the past six months. Seedance 2.5 improved on some dimensions, but does that mean 2.5 is necessarily better than 2.0?

Not quite.

Seedance 2.5 specifically boosted cinematic quality, with a big price hike to match. But on instruction following, complex motion, and other dimensions, it actually regressed.

meme bench is a short-form video generation blind test. It has 50 prompts, all hot memes from Bilibili and Douyin. We had all models generate videos from the same images and prompts, then got 400 college students and internet café regulars to vote blindly on a website — just picking which video looked better.

Our core belief: the definition of what makes a video model good should belong to the many, not to a dozen researchers.

Because this is something you can see at a glance. Generate an AI video, and who's better is obvious.

meme bench uses the Elo rating system common in chess tournaments: beating a strong model earns more points, losing to a weak one costs more, and all pairwise matchups get connected to calculate relative strength between models. This prevents rankings from being skewed by too many or too few labels on any single model.

One core goal only: SOTA models must be chosen by the people! Evaluation must serve the masses!

Very surprisingly, MiniMax H3 took first place, with scores close to Seedance 2.0 — and both slightly edged out Seedance 2.5.

We initially thought it might be insufficient sample size. After expanding test volume, H3 and Seedance 2.0's lead actually widened — H3 ultimately maintained a 55% win rate against Seedance 2.5.

Here's our testing site. You can take the test yourself, and your scores will enter the leaderboard. See which models you end up picking.

https://video-arena.com (recommended to open in browser)

Special note: meme bench accepted zero sponsorship. We also couldn't have predicted that hundreds of internet café regulars would produce this result.

Let me explain how meme bench was constructed.

The original intention was to measure video model capability in a more fun way.

So we manually curated 50 popular memes from Bilibili and Douyin, and wrote exaggerated prompts for each image to test scenarios the models likely weren't specifically trained on.

Like asking Brother Crow to spin like a helicopter.

The benefit of making model evaluation more accessible is that we can harness the power of the masses. Qiu Mu hits internet cafés more than I do, so he mobilized the café crew, and through college student gig groups and other channels, eventually recruited 400 person-sessions of labeling aka completing a full round of test questions.

The biggest challenge in college student gig groups: these folks are used to getting a red envelope for surveys, so when we tried paying them directly, we kept getting kicked out as scammers.

Before Seedance 2.5 and H3 launched, the meme bench leaderboard was very simple — Seedance 2.0 massively led all other models. Its generated short videos had the highest completion quality, best instruction following, and was obviously dominant.

For example, in the video "a man's nipples turn into titanium alloy and fire lasers that slice a house into four pieces," almost only Seedance 2.0 literally followed the instructions. Other models were too timid, with lasers firing from wrong positions.

Especially in complex scenes, Seedance 2.0's footage was noticeably smoother and richer, even self-editing and operating the camera. But the problem: Seedance 2.0 sometimes had discontinuous motion, skipping intermediate steps. In the video "a man's phone suddenly discharges electricity, electrocuting both the man and fish in the river," Seedance 2.0 jumped straight from the phone discharging to a field of dead fish, with no conduction path shown.

This is also how MiniMax H3 came from behind: extremely strong instruction following capability.

In our meme bench, MiniMax H3 was almost like an upgraded Seedance 2.0. Across multiple tests, H3 achieved coherent, smooth footage — significantly better than Seedance 2.0, which would omit processes in complex motions.

For example, in "a crowd of people crawling on the ground in twisted positions, connected into a centipede shape," Seedance 2.0 skipped the process of people connecting into a centipede, making the footage extremely bizarre.

H3 also showed better capability at generating complex scenes. In the "green screen with tons of small birds emerging" video, H3's motion was more natural than Seedance 2.0's. Especially in how the flock of birds appeared — H3 had a gradual progression, while Seedance 2.0 had a blob of birds violating physics by gushing out.

Against second-tier video models, H3's advantage was even more pronounced. Especially in intense action scenes, second-tier models were generally quite choppy.

As for Seedance 2.5 in third place — its strengths and weaknesses were both prominent.

Seedance 2.5 is the most realistic, most cinematically textured model.

In the video "a refrigerated truck door opens, and chickens, ducks, fish, and pigs chase a man," other models showed obvious AI artifacts, while Seedance 2.5 looked completely like original film footage.

But beyond color grading and cinematic texture, Seedance 2.5's other capabilities didn't significantly improve, and its motion coherence in complex scenes actually regressed. In wild scenarios testing model creativity, Seedance 2.5 performed so poorly that it made us suspect it had been specifically trained to suppress imagination.

Only when not involving complex motion did Seedance 2.5 generate more detailed, more realistic-looking scenes.

In similar scenarios, MiniMax H3's output lacked that realistic texture, more like a Bilibili shitpost video. In the video "a man gets hit by a brick but the brick shatters into powder, his head shines brightly, and 100 bodhisattvas appear behind him," Seedance 2.5 showed smooth, coherent shot transitions, with details like the bodhisattvas and blade cuts being particular strengths.

Our post-test conclusion: Seedance 2.5 is a severely lopsided model. Its most serious problem is discontinuous complex motion. For example, in "a man turns and walks to a wall, rapidly grabs bricks with one hand and passes them to the other, which launches them like cannonballs one by one, blowing up distant buildings," Seedance 2.5's output deviated severely from the prompt, with jarring jump cuts between near and far shots.

This shows Seedance 2.5 excels at generating cinematic large scenes and specific human details, but lacks coherent motion connecting the large scenes to the details. It's essentially an advanced PowerPoint.

Factoring in Seedance 2.5's pricing, this simply isn't a model for ordinary people to use casually — it's purely for studio directors.

Space is limited, so we can't break down every model individually. The full analysis report is at this Lark link:

https://likczh6fsao.feishu.cn/docx/KorydhPbIoKR8TxU2t6chhDfnGd

After H3 unfortunately became leaderboard #1, we anticipated controversy. We increased labeling volume further, trying to see if Seedance could reclaim the top spot.

H3 still maintained the highest win rate. Fat Cat is a video model that has withstood the people's choice!

The one caveat: this round tested image-to-video generation, stressing complex scenes, intense motion, and instruction following — perhaps exactly H3's comfort zone, and also where Seedance needs to catch up.

If you're unsatisfied with these results, go label on the website yourself and pick your favorite model.

meme bench has one core purpose only: beauty should not be defined by a small clique of researchers or film industry professionals.

Striving with all your might to imitate film doesn't make a better video model — it may actually run counter to popular aesthetics.

Because large models that distill the knowledge of the entire internet must serve everyone. Having a few dozen CS-background researchers define model aesthetics is a primitive and barbaric development approach. In fact, a vocational college grad's judgment of model aesthetics may be more convincing than a PhD's.

meme bench will continue updating. Any competitive new model will be added to the website — folks can come challenge the leaderboard anytime.

This is step one of our video bench series. Going forward, we'll launch more scenario-specific, more granular benches (advertising, film, gaming, memes), striving to comprehensively evaluate the most trustworthy video models for everyone.

Of course, we'll also launch a series of world model benches and AI digital human PK benches — by which I mean, look forward to Vivix showing what it's got.

(Cover image generated by ChatGPT, text purely human-written)


Subscribe to our Substack: funeralai.substack.com