Large Language Models Are a Brutal Manufacturing Business

Get Yu Hao to train the model, stat.

"Somebody Get Yu Hao to Train These Models"

"The LLM era is wrapping up; the new era of AI applications has arrived. To gear up for the next phase of AI apps—brimming with vitality and competition—Funeral AI is bringing back its old chess-move segment. This is the first installment, focused on large models. Open-source and applications will follow in the next day or two."

After two months of model labs carpet-bombing us with relentless updates, I think we've all come to realize one thing:

There's no fundamental difference between these model labs. They're all roughly at the same level. The only gap is who ships their new version a few days earlier, who ships a few days later.

GLM 5.2 was great—punched through coding scenarios with post-training. But a month later, K3 dropped and blew it out of the water.

You thought Zhipu AI was king of post-training, unstoppable at curating coding data. Turns out there was another boss: DeepSeek V4 Flash hit GLM 5.2-level performance with one-third the parameters.

You thought Kimi had cracked the code with its parameter arms race, lapping competitors by an order of magnitude with 2.8T. Then Uncle Qwen open-sourced a 2.4T giant MoE model with basically the same architecture.

From where we stand, LLMs have no secrets.

The public has a serious misconception about large models. LLMs are manufacturing, not alchemy that only child prodigies and mythical sages can master.

Manufacturing follows manufacturing rules. LLMs have no moats. Any lead is temporary—month-long at most. Other teams will catch up. Competition ebbs and flows.

Everyone's doing the same thing: sourcing compute, building data pipelines, then scaling the hell out of it. Brute-force miracles.

Any innovation gets copied by peers within a month or two.

Zhipu AI was first to crack the post-training path for coding, first to build that data pipeline, then scaled hard and detonated the market with GLM 5.2.

Every model lab does this now. Except second-tier lab Seed, which claims it doesn't distill—everyone else is guzzling Claude data.

Case in point: you can't even find a stable Claude relay service anymore. Every relay's been distilled to death by model labs. No stable Fable 5 API to be found.

Parameter bloat isn't a secret either. GLM, MiniMax—all training multi-trillion-parameter models. Kimi's lead lasted a month. Unlikely to last another.

Another case in point: data companies have become the best business in AI right now. Hundreds of startups have emerged; one already landed a major result. This is the VC market's next big theme after world models.

Markets chase hot trends, easily dazzled by some model version hitting a supposed inflection point. But the truth is: any LLM team that's made it to the table is basically at the same level.

Qwen, Kimi, Zhipu AI may lead for a moment. Hunyuan, Seed, MiniMax, StepFun catching up is only a matter of time. LongCat, Mimo have shots too—if Uncle Wang and Lei Jun throw more resources at them. Only ERNIE Bot looks truly cooked.

Because LLMs are an engineering problem. No radical architectural innovation. Scaling law keeps chugging along steadily.

I love this passage from Dario, the evil A-corp boss:

"Every few months, public sentiment either becomes convinced that AI has hit a wall, or becomes excited about some new breakthrough that will fundamentally change the game. But the reality is that, beneath the volatility and public speculation, AI's cognitive capabilities have been improving steadily and relentlessly."

Saint Liang also told everyone: don't treat them as geniuses, but as ordinary people focused on doing the work. Saint Liang speaks truth—we should listen.

Like how DeepSeek V4 Flash was unbeatable on price-performance before the price hike, setting expectations sky-high. Then the Pro official release barely improved over Flash, and tripled the price—effectively killing itself.

When Baby Whale missed expectations, Uncle Tang at Zhipu AI immediately served up a new version. GLM 5.2 base model still had headroom for post-training. Extending the "post-training king" lifespan, it is what it is.

But the difference is negligible. This pile of new models all sit between Opus 4.8 and Opus 5. Too homogeneous. Only price and speed differentiate.

Overly granular version updates aren't great either. Model labs slicing the sausage between Opus 4.8 and Opus 5—nothing but market cap management, pure attention extraction.

Good food takes time. Nobody needs to call the family until someone claims full Fable 5 surpassal. ❤️

Side note: some might say V4 Pro official release needs "minimal mode" enabled in DeepSeek Harness to show best performance.

But look at all those qualifiers before "official release"—isn't that a bit funny?

How's this different from car manufacturers swapping every component and hiring pro teams to tune for Nürburgring lap times? I think piling qualifiers before benchmarks is a bad habit spreading from the auto industry to LLMs.

Though LLMs and EVs really are alike: no miracles, just methodical engineering capability.

Only three variables determine LLM outcomes: compute, data, and organizational capability.

Every model team at the table satisfies at least two. If compute and data are fine but model capability still lags, it's definitely an organizational problem.

The textbook case: Gemini.

Turns out scientists can't lead LLM teams at this stage. Scientists don't like managing labeling crews, don't like engineering grunt work. But right now, LLMs need foremen driving 007 shifts with the team—main job is problem-solving, test-taking.

And that, incidentally, is what Chinese model teams are good at.

So you see, China's model teams are all decent. Top teams take turns competing for world #3. Second-tier teams like Tencent, Meituan, Xiaomi can all ship a not-unusable model within six months.

A hard indicator for whether an industry is manufacturing: can Elon Musk figure it out?

If Saint Ma crosses over and can't crack it, it's definitely not manufacturing. E.g., Saint Ma doing politics—total sideshow, getting walked by Donnie like a dog.

Conversely: rockets, Tesla, Grok—Saint Ma didn't get them at first, then did. The inflection point was when the industry's infrastructure matured enough to become manufacturing. Just do it.

I can't help but wonder: has there ever been a worse business than LLMs in human history?

Tech giants pouring hundreds of billions in capex, begging you, pleading with you, giving you their latest products for free.

The model teams themselves are burning through lives too. High pay is one thing, but they might not get to spend it. Every model worker I've seen looks haggard, dark circles, acne, greasy hair for women, hair loss for men.

The root of brutal LLM competition is homogenization.

All models are competing on coding and Agent scenarios, chasing marginal gains on SWE and a few other benchmarks. So the public has only two choices: use the best, or use the cheapest.

But models can differentiate.

Our beloved Kimi K3 is a model. You don't just need to recognize frontend scenarios' importance—you need to build data pipelines, precisely define data for one-shot generation of polished web mini-games, then scale hard.

That's product definition capability. Everyone should study it.

At this stage, model productization is massively undervalued. Because researchers training models don't have product definition skills. Defining products is creative work—it doesn't scale.

Of course, this kind of product definition isn't unique to Kimi. Our erratic-but-brilliant MiniMax gets it too.

H3 video model heavily optimized for film post-production capabilities while beefing up prompt engineering. This made H3 better than Seedance 2.5 at generating fun short videos.

Specifically, our meme bench, labeled by 400+ internet café family members, ended with H3 scoring above Seedance.

After meme bench went public, thousands more Funeral AI reader family members came to label. Given Funeral AI readers are all highly educated, high-quality, high-culture "three highs" people, H3's lead actually widened—proving H3 is genuinely better at generating good meme short videos.

Side note: meme bench (video-arena.com) just added Wan 3 model. Welcome family members to come label and see what Wan 3 is actually made of.

Of course, we can't say Seedance 2.5 lacks capability. Rather, 2.5 deliberately thinned prompt engineering. Requiring users to write detailed second-by-second prompts actually improves rigid instruction-following, better suited for professional creative scenarios.

That's a model product definition choice.

LLMs really are manufacturing. The exact same thing happening in solar, small appliances, and new energy is happening here.

Take hair dryers. You can't compete with Dyson by just making hair dryers. But Dreame's Yu Hao developed ultra-high-speed motors and redefined the product.

Wei Xiaoli—nobody can say who's technically ahead. But NIO has battery swap. Li Auto has its relentlessly upgraded fridge-TV-sofa. In industry homogenization, product definition creates differentiation.

So right now is actually a great time for applications.

If a data company came out today and said, "K3's frontend excellence came from our mini-game dataset"—that company would get mobbed by thirsty model labs.

Of course this is hypothetical. Without the great Yutong Zhang and mysterious advisors, I don't believe Kimi could've built these differentiated capabilities.

But the Cursor example is right there. Today's application companies actually have better model-defining capabilities than researchers training models.

So model productization isn't a job title—it's an ecological niche for a category of application data companies.

Application companies have product definition capability and real user data. They can make model capability improvements serve real users, not benchmark-gaming vanity.

That's also why Funeral AI co-publishes Bench with application companies.

All these principles, actually, were taught many times by the greatest entrepreneur in my heart, Teacher Yu Hao.

Sadly, Yu Hao stopped his daily hundred-short-video streak after some setbacks. I haven't heard him tell the story of Tsinghua tech genius developing ultra-high-speed motors to remake hair dryers in three months. But Yu Hao's breakthrough thinking is worth studying.

Honestly, Yu Hao—a heaven-sent manufacturing giant—is perfect for LLMs.

But Kimi's doing well now, and there are too many gods. Probably doesn't need Yu Hao. But MiniMax, still proving itself with something big brewing—MiniMax needs him.

I formally recommend MiniMax hire Yu Hao to train models. First bring Yu Hao in to define data. Dreame's R&D methodology—all of it goes into model training. When the new model launches, bring Yu Hao back to shoot short videos connecting hundred-thousand-RPM high-speed motors with 3T-parameter LLMs.

Double Win. Win-win!

(Cover image generated by ChatGPT. Purely human-written content.)

Final announcement:

Welcome to the first "Model Emperor" Bench Hackathon, co-hosted by Funeral AI and Qwen—putting various models to the real-world test offline.

Qwen 3.8 Max, K3, DeepSeek V4 Pro, Opus 5, 5.6 Sol APIs supplied in full. Use your real work scenarios to interrogate models—what works for you is #1!

Scan the poster QR code to register. Registration confirmation emails go out Wednesday noon.