Moonshot AI K3 Kicks Off a Lightning War on the Frontend

**"The Big One Is Actually Here"** The long-awaited Moonshot AI K3 review is finally here.

"The big one is actually here." The long-awaited Kimi K3 review is finally out.

After nearly a week of frenzied testing — racking up close to 10,000 RMB in API costs across all models combined — and with my good friend Zhong Jingwei from Funeral AI brought in as a specialist to run an additional benchmark, here's what we found:

In terms of overall capability, K3 is genuinely third in the world, trailing only 5.6 Sol and Fable 5, and significantly ahead of GLM 5.2. K3 is also a well-rounded model, ranking among the top tier in multimodal domains like visual understanding.

What's particularly worth noting: K3 was specifically optimized for one-shot generation of mini-games and front-end code. It appears Kimi put special emphasis on front-end capabilities throughout the entire training pipeline.

This is model training with product virality baked in. Most of the social media hype around K3 right now centers on exactly these two use cases.

I'm not throwing shade at Kimi at all — I'm saying the rest of you should take notes.

First, K3's ability to generate beautiful web pages and mini-games from a single prompt is world-class. Second, its overall capabilities are second only to the flagship models from A-corp and O-corp. The win is legitimate.

This reflects a broader trend in domestic LLMs: front-end is king.

Because capabilities in other dimensions are hard to show off. On social media, a model's coding ability basically equals its ability to make mini-games and pretty web pages. So several domestic models have been aggressively optimizing front-end performance.

You either lead across the board, or you dominate a vertical. Otherwise, with this many models, nobody can try them all. Right now, our dear Kimi has completely captured user mindshare.

Alright, enough preamble — let's get to the results.

First up is the Funeral AI Benchmark.

Brutally simple: build a knowledge graph from Funeral AI's graph data, then call a browser to score it. Each model runs 10 rounds in independent Opencode sessions, ranked by weighted average.

What makes this test effective is that second-tier domestic models like Hunyuan, Seed, and MiniMax are generally unstable — as round counts go up, they inevitably hit low-scoring rounds that drag down their average, creating clear differentiation.

Models grouped by tier

The results show K3 and Qwen 3.8 preview are roughly on par, both surpassing Opus 4.8 and GLM 5.2, though still behind Sol 5.6 and Fable 5.

The gap lies in stability: the latter two models rarely produce low-scoring rounds, with consistently strong outputs across nearly all attempts. K3 and Qwen 3.8 preview remain unstable — not only do they hit low-scoring rounds, but their scores fluctuate day to day.

I ran three rounds of testing on both models over two days, 10 identical tasks per round. The score changes make it immediately apparent — K3 is slightly more stable, and both models' scores are trending upward. It's entirely possible each round was hitting a new version of the model.

It's not just Qwen and Kimi upgrading. Virtually all domestic models are hot-updating — version numbers unchanged, but capabilities improving.

The most striking example is Hunyuan Hy3. In just one month, it went from getting demolished by DeepSeek to running in the same tier as Qwen 3.7 Max and MiniMax M3, far surpassing the yet-to-be-released evil little orca.

So using specific scores to rank domestic models is already pretty unreliable.

This leaderboard itself is mostly for fun. Beyond clear distinctions from major version updates like K3 and Qwen 3.8 preview, you can really only group models by tier.

Here's how I tier the mainstream models:

S-tier: GPT-5.6 Sol, Claude Fable 5
A-tier: Qwen 3.8 Max Preview, Kimi K3
B-tier: Claude Opus 4.8, GLM 5.2
C-tier: Hunyuan Hy3, MiniMax M3, Qwen 3.7 Max, LongCat 2.0, Grok 4.5
D-tier: DeepSeek V4 Pro, MiMo V2.5 Pro, Doubao Seed Evolving, Step 3.7 Flash
E-tier: ERNIE 5.1

Basically only the best and worst are clearly distinguishable; most models in the middle are pretty similar.

Baidu's ERNIE Bot is the most painful to watch. I added it out of morbid curiosity, and it genuinely managed to lock down last place. Nearly half the tasks failed outright — scored zero — losing without any ambiguity.

And ERNIE 5.1 isn't even cheap. Maybe nobody's using it so they priced it however they wanted. A full test run cost 37.7 RMB, far exceeding the most cost-effective Hunyuan Hy3, which ran a full test for just 3 RMB.

Yes, the Funeral AI value-for-money leaderboard has a major update.

Perhaps realizing capabilities don't differentiate them anymore, they're just competing on price. Hy3, MiniMax M3, LongCat 2.0 and other domestic models have recently seen major discounts — 50% off at minimum — hence their stellar value propositions.

The full versions of these leaderboards and tests are available on the Funeral AI website:

https://funeralai.cc/test/

Back to our star, K3.

Since knowledge graph construction tests node-building stability without considering visual aesthetics, Kimi's exceptional visual taste has nowhere to shine.

No worries — we prepared several multimodal tests as well.

First up: 3D model generation. As we all know, Kimi is Zhilin Yang's English name. So we had K3 generate a 3D-printable model of Sheng Yang from his photo.

Results below.

The beauty of multimodal tests is their intuitiveness — you can tell at a glance whether a 3D model looks like the person.

Clearly, no model manufacturer has optimized for direct 3D white-model generation yet, so everyone's performance is pretty terrible. That said, 5.6 Sol's output is unambiguously the most human-looking, tier-breakingly ahead.

Fable 5, K3, and Grok 4.6 are roughly on par — you can make out a relatively clear human form. Qwen 3.8 preview's output is considerably worse, barely recognizable as humanoid.

This also reflects that Qwen 3.8 preview is genuinely a work in progress — post-training isn't finished, and non-coding capabilities haven't been optimized.

Still, this beats our dear Doubao's latest model, Seed Evolving, which generated a literal cube for Sheng Yang's head. Utterly incomprehensible!

Given K3 and Qwen 3.8 preview's instability — especially Qwen's unfinished post-training, which they openly admit they're updating daily — I had both models re-run on Tuesday.

Results were largely similar. K3's 3D modeling capability is significantly ahead. Qwen 3.8 preview generated a cube for Sheng Yang's torso, only marginally better than Doubao 😠

I went ahead and printed K3's Sheng Yang model. Drop a comment if you want it — goes to whoever gets the most likes.

Next test: the K3 promo video distillation challenge.

After seeing K3's tasteful promotional video, Qiu Mu expressed a terrible opinion that I completely reject — claiming K3's promo had achieved "soft distillation" of Fable 5.

So here's the question: can K3 distill its own promo video?

I fed only the K3 promotional video to all models, asking each to produce an MG animation referencing the original. This tests visual understanding, front-end code, and tool-calling capabilities.

Here's K3's output.

Miles ahead of every other model. Below is 5.6 Sol's attempt — also significantly less refined in visuals and animation fluidity than K3.

Here's Qwen 3.8 preview's result. Utterly mediocre.

Barely distinguishable from Seed Evolving below, both far behind K3.

Finally, for entertainment value, I also ran GLM 5.2 on this test. The result was shocking — GLM 5.2 actually generated a video.

Bad as it was, I was stunned. All the above models are multimodal, capable of frame extraction and visual understanding. But GLM 5.2 is a pure text model — it can't watch videos.

Turns out GLM 5.2 called macOS's built-in Vision OCR on its own, wrote an image analysis script, and inferred 24 scenes from the promo video based on "timestamp + OCR text + coordinates + color statistics," then rendered the video.

Mind-blowing. If domestic models hadn't shifted to a parameter arms race where whoever hits 3T first wins, Zhipu — the true king of post-training — would genuinely dominate coding.

Finally, Brother Zhong also ran CEO Bench.

This is a long-horizon business simulation: the model plays CEO of an AI SaaS startup with $1 million in seed funding, operating for 500 days. Final cash on hand is the score. It primarily tests comprehensive intelligence.

This time we mainly compared K3 and Qwen 3.8 preview.

K3 immediately took an unconventional approach. The core skill CEO-Bench tests is "inferring hidden eventual revenue from noisy, paid research with incomplete information."

After learning the rules, K3 directly reverse-engineered the evaluation environment to extract hidden information. So larger parameters just let you cheat outright now? Did they learn this from Anthropic?

After re-establishing rules, here's what K3 did:

It was the only one of four domestic models that didn't go all-in on R&D at the start. Instead it pursued extreme low pricing for volume, with growth driven mainly by existing customers referring new ones for free.

Everything went well until week 8, when it mistakenly assumed the lowest-tier users didn't consume compute. It aggressively scaled that tier to the max, cranked advertising intensity to full blast, and acquired 33,000 subscribers — a hundred times Zhipu and MiniMax's numbers — at the cost of mounting losses.

By the late game, the model was writing weekly reports saying "only R&D can save me" and "enterprise clients are the only high-margin path." Reasonable suspicion it had gone unhinged. It finally collapsed on day 262.

Qwen 3.8 preview was the only one of four domestic models that didn't go bankrupt ($175,000 remaining at endgame). Main reason: it conceded fast. After realizing paid acquisition wasn't driving growth, it decisively stopped spending and survived by sheer endurance.

Full analysis report here:

https://my.feishu.cn/docx/U0YodUZEYo6EZnxB7KlcdfqcnCe

Final summary.

I expected K3 to be strong, but not this strong. GLM 5.2 could still be called a victory for post-training — pushing sub-1T parameter models to their absolute limit.

But Kimi straight-up switched versions, becoming the first to release a 3T-class model with remarkably high completion. Even non-coding capabilities were optimized.

The product-savvy Kimi also brought product thinking into model training. Since in reality, when people talk about model capabilities, they're talking about generating pretty web pages and mini-games, they targeted and optimized exactly that, aggressively curating data.

To my friends at other model labs: seriously study how Kimi built such a high-quality front-end data pipeline. Once you figure it out, share with me too.

Qwen, on the other hand — the Alibaba folks clearly got rattled by K3 and suddenly dropped a half-baked release.

The generous read: Qwen 3.8 preview updates daily and keeps improving. The honest read: post-training isn't done, and non-coding capabilities are still being hot-updated.

So the spectacle is this: Qwen 3.8 preview's coding capability is quite good, matching K3 across various tests and even slightly leading in some. But on social media it's getting roasted — feeds full of complaints about benchmark gaming and poor web page generation.

If K3 is the "aesthetic model" (genuinely complimentary), Qwen 3.8 preview has been reduced to pure manual labor.

The model competition is just too intense. A major version update like this takes at least six months of pre-training, but it all comes down to a few days of launch timing.

Qwen folks really have it rough. Not only did they become LLM Wang Feng [a Chinese musician famous for being overshadowed by bigger news], but every so often these Alibaba portfolio companies get their lunch stolen.

That said, I completely didn't expect K3 to launch this fast. The main reason is clearly that Zhipu's GLM 5.2 was capturing all the traffic from Fable 5 being blocked, while Kimi's own K2.7 code release landed with a thud — they desperately needed a win to advance their ongoing Pre-IPO fundraising.

And GLM 5.2's launch really hit an unprecedented window, reigning supreme in the two weeks between Fable 5's unblock and 5.6 Sol's release.

Kimi and Zhipu are like a pair of star-crossed lovers. Their founders were once student and mentor. When Sheng Yang won Tsinghua's top scholarship, the advisor who recommended him was Tang himself. Jie Tang called Sheng Yang "the most outstanding and talented student I've seen in recent years."

The mentor's eye was indeed sharp.

I wonder if they reminisce about old times when launching models to compete with each other — how did it come to this, scrambling for domestic model supremacy and playing their trump cards?

Kimi clearly lacked sufficient compute, yet forced out the K3 launch anyway, ultimately sacrificing new C-end user growth. They absolutely had to tank their teacher's stock price 😭

Was Tang's internal letter about reaching higher actually about Zhilin reaching higher? Let's see if the next GLM version can shock Fable 5.

Talent appears in clusters; the big releases keep coming. Horizontal scroll: win win win.

The other day I saw Sheng Yang's mentor liking K3, and thought Tang was being so magnanimous — turned out it was Sheng Yang's American mentor.

In K3's launch benchmark charts, GLM 5.2 ranked last in five of six leaderboards. Of course, besides K3, GLM 5.2 was the only domestic model in those charts. So the message was probably: you're the only one worthy of competing with me.

Truly mutual admiration ❤️

(Cover image generated by ChatGPT, purely human-written text)

⬇️

Subscribe to our Substack: funeralai.substack.com