Shengshu Technology Diagnosed With Duck Leg Overdose

If you had to pick the most orthodox Tsinghua-affiliated AI startup, Shengshu Technology would be it: all three founders are from Tsinghua, all from the same research group, still physically on campus — essentially a university lab doing corporate tech transfer. Gripped by my own pathological worship of meritocracy, I dutifully tried their product Vidu, and came away thinking: folks, we've had one too many goose legs. Because looking at their model capabilities, product positioning, and marketing strategy, the student energy here is absolutely off the charts — it's completely maxed out.

You could say Shengshu Technology is the most orthodox Tsinghua-affiliated AI startup: all three founders are Tsinghua grads, all from the same research group, operating right on campus — basically a university lab doing corporate tech transfer.

Fueled by a pathological reverence for meritocracy, I reverently tried their product Vidu, and concluded: folks, we've had too many goose legs.

Because from model capabilities to product positioning to marketing strategy, this company reeks of student energy — it's completely saturated, on the verge of emergent intelligence.

First off, I never expected that in June 2026, opening Vidu's website would reveal their brand new feature: ViduClaw.

And they're still updating this documentation.

The last true believer in crawfish in North China, the final soldier of OpenClaw.

Even that crowd who used to tour the country running crawfish qigong workshops has packed it in, and you're still here like "crawfish, my crawfish."

And Vidu genuinely wants to make ViduClaw the entry point. No matter what page I click into, some popup jumps out to remind me: come try our latest crawfish feature!

Hard to refuse such hospitality. So me, who's never even deployed OpenClaw, had to nervously try out this trendy feature.

Turns out it's fine, just a chat box, no deployment needed.

But I don't understand the point of this thing. In Vidu's documentation, they describe the difference between ViduClaw and other model products like this:

Natural language dialogue, one-sentence generation... buddy, do you think it's still the Stable Diffusion era? Which product on the market can't generate video from natural language now? Which video model doesn't have built-in Agent mode?

Instead, Vidu took functions that other model products package up nicely and unpacked them, turning them into Skills uploaded to GitHub for users to download as needed.

I feel like the only audience for this workflow is overachieving grind students.

Be like⬇️

All of it, just to flex on us dullards with one question: pretty hard, right?

Well you win, I lose, my mediocre brain genuinely can't figure out elite crawfish, I can only handle foolproof products that hold my hand the whole way 😭

Anyway, I still tried it, wanted ViduClaw to batch-produce a few episodes of a short drama about Tsinghua's Goose Leg Auntie for me.

It asked me for prompts, had me fill out a form.

Really treating me like a student, itself like a counselor.

No choice, I had Claude generate prompts and sent them over. Result:

Yep, no matter how I modified the prompts, ViduClaw insisted this was prohibited content and couldn't generate it.

Safety compliance isn't the issue. But I ran the exact same prompts through Dreamina, Keling AI, and PixVerse without any problems.

The funniest part is, when I stopped using ViduClaw and switched to regular text-to-video, the video generated just fine.

So your crawfish is a safety auditor?

The story I gave Vidu: A white-collar worker in Beijing's CBD holds up a phone with nothing but green on screen and asks Goose Leg Auntie why his stocks are green, and Auntie says it's from a harmless green new energy juice marinade.

Authorized

I batch-generated 3 videos, and the most watchable result is below:

Really disappointing.

After all, Vidu claims to be "born for drama" with "synchronized audio and visuals," so I had high hopes for characters speaking lines with lip-sync.

But in every generated video, the lip movements don't match, the lines don't match the characters, and they frequently spout nonsense.

As for character movements, object stability, background details — there's no point even evaluating those.

Hard to believe this is a 2026 video model. The whole thing has this real-time world model generation feel — fuzzy, chaotic.

Might as well follow PixVerse's lead and pivot to world models, switch tracks so everyone feels better.

Oh wait, Shengshu's "Motubrain, the first general world action model to top two authoritative embodied AI benchmarks" is already on the way. We'll see.

I also wondered if maybe the plot was too hard, or if the prompts themselves were the problem. So I tried PixVerse and Dreamina.

PixVerse:

Actually the lip-sync is mediocre too, but at least the right person is talking, and the movements are passable.

Dreamina:

Has that AI feel, can't call it perfect, but at least no issues with movements and lip-sync. Better than both Vidu and PixVerse.

I feel like you need to reach at least this level before claiming to be "born for (AI short) drama."

So where does Vidu's confidence in being "born for drama" come from?

After careful consideration, I think it's probably video length.

Current mainstream video models, whether Dreamina, Keling AI, or PixVerse, generally cap single-generation videos at 15 seconds.

Our Vidu? Insists on 16 seconds, towering above the rest by a single second.

Feels like they want to hire Yue Yunpeng as spokesperson and belt out: Ah, 16 seconds, one second more than 15. Product promo directed by Zhang Yimou, titled One Second.

Of course, back in 2024 when Shengshu published that Vidu paper, those 16 seconds made history.

After all, Sora was still just an internal demo, Dreamina and Keling AI didn't exist yet, and those college-kid projects competing for domestic video model supremacy couldn't generate more than ten seconds. Vidu announcing 16 seconds of continuous generation was instantly deified.

But a launch is just a launch, a paper is just a paper. By the time Shengshu officially released Vidu Q3 with 16-second support, it was January 30, 2026 — and just days later, the glorious Seedance 2.0 went live, making Vidu's 16 seconds purely for self-consumption.

Pathetic. The videos are unwatchable, only the duration wins by a second.

Vidu is exactly like that grind friend you haven't heard from in ages: pulling all-nighters doing practice exams in high school, grinding GPA and comprehensive assessments in college, no internships, no socialization, finally showing up to Big Tech interviews with a resume covered in student council titles and teacher recommendations, landing zero offers.

So at the reunion, taking a sip of baijiu, shedding tears and sighing: Ah, maybe I'm not at the table now, but back during the Hundred Model War, I had several more seconds than you all, remember that paper...

Shengshu is living in the past, but in the AI era, model iteration is fucking fast, competition is fucking fierce. Its former domestic rivals have pivoted or exited, and those remaining are basically Dreamina and Keling AI sheltered by Big Tech, plus AISphere's PixVerse.

Pretty far from ByteDance and Kuaishou now, so Shengshu can only 1v1 AISphere, mentally replaying the intensity of yesteryear.

But both companies' video generation capabilities belong to the second tier, so they can only compete in storytelling contests — telling commercialization stories, and also stories about things beyond AI video that can't be commercialized.

PixVerse does world models, Vidu does embodied brains, each chasing the other, both terrified of dropping to the third tier and sharing a table with the wrapper crowd.

The latest story: Shengshu is launching a Hong Kong IPO, and AISphere quickly followed with its own listing rumors. Even this is a race.

Please, both of you, spare some energy for actual AI video 😭

That said, I do think Shengshu has a genuine fixation on video length.

Because one of ViduClaw's key features mentioned above is automatically stitching multiple short videos together into one long video. Theoretically unlimited length.

To demonstrate this, I opened ViduClaw again and gave it a task: make a crossover animation combining Fat Cat and Goose Leg Auntie.

Plot as above

I even thoughtfully uploaded 3 reference images, asking it to generate video based on these.

I sent the request around 7 PM, it dawdled along, and didn't finish the video until around 9-10 PM.

And the resulting 30-second video looks like this⬇️

Despite my uploaded reference images, every Fat Cat frame has inconsistent art style — you'd think this was some deliberate artistic choice.

And the final Goose Leg Auntie image uses Fat Cat's too. This crawfish is deaf and blind.

What does this have to do with Vidu's promoted "subject consistency"? Zero subjectivity.

Anyway, everything about Vidu is student-energy, the generated videos feel like group project work.

Including opening their WeChat official account — the content style is exactly like those university official accounts.

Just this whole AI-circle banquet-table vibe, anyone feel me?

And recently, realizing its video quality can't keep up, Vidu started going the cost-performance route, mainly targeting middle-aged B2B bosses.

"Price slashed 20%! Speed surged 20%! Most cost-effective video model Vidu Q3 is here!"

Not lying. My rough calculation: the money for 1 Dreamina video gets you 3 Vidu videos.

But of those 3 Vidu videos, not a single one is usable.

Especially with Seedance mini coming too — if that hits at 30% discount, doesn't Vidu lose its niche entirely?

So Vidu's product thinking is pure student mindset, thinking if they just work harder and harder, like manual laborers grunting away, working for free and sacrificing everything, clients will be happy and pay up.

In reality, spending ¥3000 to hire three college kids bumbling around for a month is worse than hiring one expert for a day.

Vidu really needs to get Feng Ge to give them a lesson, staying in the ivory tower like this is finished.

Of course, since launching Q3 this January, Vidu hasn't had any major updates.

Whether they're cooking up something big, hard to say. I'd genuinely love for Shengshu to release a truly "born for drama" video model that slaps my face swollen.

But until then, eat fewer duck legs, lose some of that student energy, and get out into society more.

(Cover image generated by ChatGPT, purely human-written)