MiniMax Steals Big Results with H3
"Saving the model through a detour" — a pun on 曲线救国 ("saving the nation through a detour," a historical euphemism for indirect approaches), applied to AI model development or distribution strategies that work around direct obstacles (regulatory, technical, or competitive).

"Saving the Model Through Indirection"
Last weekend MiniMax dropped H3, and suddenly the internet was flooded with reviews claiming they'd found a budget alternative to Seedance 2.0.
This take is funny for three reasons:
First, sure, it's "budget" — the price is one-third of Seedance 2.0. But since when is "cheap" a selling point in the video model wars? Some no-name fourth-tier video models are even cheaper. You see them getting any traction?
As I always say: if creators were purely price-shopping, they might as well go all-in on voiceover PPT slideshows. Video model companies could just stay home and pray for Gresham's Law to run its course.
Second, after actually using it, I'd say MiniMax H3 still trails Seedance 2.0 by roughly 20%.
For instance, I had it generate a video of "Feng Ge apologizing to everyone who bought tech stocks on his recommendation" ⬇️
You can clearly see Feng Ge's face looks like it's caked in foundation — the texture doesn't reach photorealism, more Madame Tussauds territory. His famously wise eyes are locked in a wide-eyed stare, gaze completely frozen, like an Eastern Mona Lisa with serious uncanny valley energy. And Feng Ge and the Nasdaq exchange background might as well be on separate layers — zero integration.
Gaps like these can't be papered over with "but it's cheaper": video models follow a reverse bucket principle. If H3's capability ceiling is 80% of SOTA, no amount of regeneration gets you to 100%.
Of course, for some people 80% is enough. Nothing I can say to that.
Third — and this is the funniest — MiniMax H3 drops, and Seedance 2.5 immediately comes roaring out.
Be like: you were just trying to quietly hype yourself up, and now you gotta fight the final boss? Bro wasn't ready.
Though at least H3 launched slightly earlier, so there's that small mercy.
Anyway, I emergency-tested H3, generated 200+ videos, and here's my conclusion:
The one thing MiniMax H3 genuinely does better than other video models is actually the least "AI video" thing about it — its relationship to the real world.
Simply put: if other AI models are doing VR (virtual reality), H3 is doing AR (augmented reality).
When users ask H3 to generate an AI video from scratch, the results are underwhelming.
MiniMax seems to have figured this out themselves, so they pivoted: when users ask H3 to modify existing videos, it can actually surprise you.
This is saving the model through indirection.
A few examples:
01
Summer of Yujiro: The Release Cut
First, the greatest Chinese youth film of all, Summer of Yujiro. I had H3 transplant this classic clip into a "running like a J-drama" effects video by creator @原来是陶阿狗君 ⬇️
Left: original video; Right: final product
Pretty uncanny. I never opened CapCut or any other post-production software, yet the output looks like I applied a template. H3 isn't just face-swapping the reference video — it seems to genuinely understand the editing logic. And Fan Xiaoqin's image, movements, and voice barely change. H3 knows restraint: it's not diffusing everything from scratch, but knows what to preserve and what to rebuild. Real agency there.
This characteristic makes MiniMax H3 perfect for positioning itself as post-production software.
Right? Maybe nobody will generate video with MiniMax, but everyone will run their generated videos through it for effects, packaging, whatever — it's cheap anyway.
While everyone else is racing to turn video models into AI Agents, it just made itself a pure tool. Carved out its own ecological niche, stopped competing on the same stage as the AI crowd.
Gotta admit MiniMax has stolen some real short-term results here — it's literally the only player in this lane right now.
02
"Chinese People Can Fly" MV
Beyond effects, I wanted to test H3's text rendering. So I had it make an MV for the recently viral "Chinese People Can Fly" ⬇️
Gotta say, the effect is genuinely good.
When text floats in space, the fonts are harmonious, the gloss animations aren't tacky, and per my request it stays behind Lan Lao the whole time without blocking that genius face.
When text sticks to walls, it's not just rigidly plastered on — it follows spatial structure and physics.
I'm half-convinced they secretly embedded CapCut inside and maintain an in-house subtitle team that pauses AI service for text generation requests, just letting some intern on 3,000 RMB/month add them manually.
Not impossible — there's precedent. Any interns want to submit an exposé on whether you're actually digital laborers for 置身稀内? 😭
03
Sun Xiaochuan Meme 2.0
Next, testing H3's audio editing capabilities, since MiniMax has its own audio model.
I fed H3 Sun Xiaochuan's classic "summoning the property manager" clip, with instructions to convert his speech to Sichuan dialect, add Japanese subtitles with equivalent meaning in typical Japanese variety show "flower text" packaging, use high-frequency Sichuan opera face-changing to block Sun's face, and bleep all profanity for a double mosaic effect.
Final product ⬇️
My Sichuanese friend Xianyu's review:

"Feel like his manners are worse than his anti-fans now"
Being able to auto-translate speech into Sichuan dialect and Japanese, plus precisely sanitize the video — H3 does demonstrate some understanding of uploaded materials and the real world.
Many creators say H3's advantage is AI video post-production, but I think that's missing the point. AI videos wanting effects can just regenerate from scratch — no need for a dedicated post step.
MiniMax H3's real advantage is AI post-production for shot-on-camera video, because after its modifications, the footage still looks shot. That's genuinely rare.
04
Opening Obsessed as a Galge
Though it's post-production software, the 15-second video limit makes breaking into pure film/TV industry tough. So MiniMax fits better for short-form content: memes, remixes, packaging.
The use case I imagine: when the girlies and bros record a vlog, no more dry posting — just H3 that thing up.
I didn't have vlog material handy, so I took a rom-com-vlog-like scene from Obsessed and packaged it as a galge ⬇️
Top: original; Bottom: final product
Flaws exist. H3 straight-up omitted the female lead standing up — Nikki's ass is glued to that chair. And maybe because two-person dialogue complicates text box handling, I had to generate this one many many many times, and the final output still didn't match film dialogue with frame-level precision.
But the vibe is there — real Doki Doki Literature Club energy. Suggest all heterosexual lifestyle bloggers adopt this format going forward.
05
Reshooting The Backrooms
And H3's editing as post-production software isn't just slapping stickers on video surfaces — it treats the original footage as a space to understand.
So you can request elements added to specific locations, or perspective shifts: first-person to third-person, eye-level to overhead, changing camera movement paths, etc.
First I tried adding Nailong and Pangmao to The Backrooms — these are MiniMax's foundational IP, can't forget your roots ⬇️
Top: original; Bottom: final product
Characters and environment stay largely consistent; Nailong and Pangmao both successfully checked into the Backrooms. And both cast shadows matching Backrooms lighting, proving they're physical entities not pure ghosts 👍
Decent execution.
Then I tested perspective shifts, requesting H3 modify three shots as follows: Shot 1 to first-person; Shot 2 to dolly zoom; Shot 3 to drone shot taking off from behind and tilting down; Shot 4 to pull-back.
Results ⬇️
Top: original; Bottom: final product
Lots of problems: the protagonist's POV doesn't fully match the original, shadow placement is off, and the dolly zoom wouldn't pass Hitchcock's own standards.
Well, this task is pretty hard, and I haven't tested other models on it, so understandable.
But the perspectives are right, the camera movements are right — if you're not frame-matching against the original, the footage works.
So fam, stick to H3 for memes — adding Pangmao and Nailong and such. Perspective shifts at film industry level? Wait for another iteration 🥵
06
If Tarantino Shot a Moutai Ad
I mentioned H3 only generates properly with reference material. But "reference" doesn't have to mean uploaded video — pure text works too, since obviously MiniMax stealth-trained on tons of material.
So I tested ultra-short prompts like:
"If Nolan shot The Backrooms"
"If Tarantino shot a Moutai ad"
"If Junji Ito shot a Spring Festival Gala skit"
...
The best performer was this Tarantino Moutai ad ⬇️
Mainly because nowadays mentioning "director's skills" automatically means Wes Anderson, which is boring as hell. Per internet logic: symmetrical? Wes Anderson. Saturated colors? Wes Anderson. That style is too easy to mimic.
MiniMax H3 may have trained specifically to avoid being called out, so when doing "one-sentence movie generation," it shows understanding of different directors' content logic and narrative habits.
This Tarantino Moutai ad doesn't have especially distinctive color grading, no blood, no feet — yet somehow it just reads Tarantino.
Suggest Kweichow Moutai adopt this video directly; contact available via DM.
What I want to say finally is: MiniMax H3 making itself into post-production software is distinctive, commercializable, but feels like voluntarily moving from the big table to the small table.
Like: "You first-tier video models keep fighting, I'm already outside the three realms and beyond the five elements, focusing on post-production software, don't @ me."
Actively finding a corner position so being below SOTA becomes unassailable.
Nothing wrong with it, but MiniMax living like this feels kinda pointless.
Let me show MiniMax a better path:
Remember when Sora launched, making videos that didn't understand physics yet calling itself a world simulator. MiniMax H3 now wildly modifies shot footage, doing work that's mainly about understanding and transforming the world — might as well call itself a world model.
So many claim to be world models now, each with a meter-long string of qualifiers. H3 could launch as "the first non-Liu-style-generated · non-physics-related · non-embodied-industry · non-autoregressive · non-overseas-Chinese-team · world model." Very 带派.
Plus it open-sources like many world models. Speaking of open source, I'm reminded of LibTV — heard Mian Shen even posted a Moments supporting H3, understandable since H3 basically reattached a leg for him, seamlessly connecting to post-training video models to jointly resist Seedance 2.5.
Damn, H3 as both post-production software and open source — why's this playbook so familiar?
Fam, this is pure Adobe's tolerate-piracy strategy.
Right? I deeply suspect H3's calculation is: open-source to get all of China using it, build dependency through habit, then drop a major update to harvest beautifully — poaching everyone else's traffic in the process.
So that's what "open-source traffic hijacking" means.
All I can say is: absolutely diabolical. Please use MiniMax H3 with vigilance and critical spirit.
(Cover image generated by ChatGPT; article purely human-written)
⬇️
Subscribe to our Substack: funeralai.substack.com