Moonshot AI's Return to Glory
Good fortune lies ahead.
"Good Fortune Lies Ahead"
With OpenAI unable to push out GPT-6 anytime soon, the LLM leaderboard has been twisted into a pretzel by Chinese teams.
Aside from the top two spots permanently occupied by Fable 5 and 5.6 Sol, ranks three through ten are updating faster than HEYTEA drops new drinks.
The trend these past two months: at first Opus 4.8 was being farmed like a boss monster, then after GLM and DeepSeek dropped updates these past couple days, Opus 5 got farmed too.
But looking at specific benchmark scores is already meaningless. With so many benchmarks, so many runs, and everyone testing on their own Harness, the actual numbers depend entirely on whether a team feels like launching a hype satellite.
What I mean is, model labs can produce whatever scores they want. So public benchmarks don't matter—only real-world performance in your actual workflow matters. That's why we're building Dongmodi, pulling together personalized evaluation sets ❤️
Back to the brutal LLM race.
The situation is clear: the new season has begun.
The first half of 2026: blitz the coding scenario. Under 1T parameters, crank up coding capability through post-training. The season's undisputed king is clearly Zhipu AI, whose GLM 5.2 completely broke through user mindshare, first to slay Opus 4.8—first through the pass, first to rule.
Second half: parameter leapfrog. Whoever ships a stable, usable 2T parameter model first wins. Post-training is endless, but pre-training hype satellites reach even higher.
Season just started, and Kimi K3 was first to launch the satellite—a child of the patch. Its scores approach Fable 5, switching the domestic model narrative from "budget alternative" to "challenging for #1," and with extreme optimization for one-shot beautiful webpage generation, it broke through the frontend scenario, kicking off a new season of domestic model moonshots.
K3's frontend one-and-done capability is genuinely impressive. When I built the Mianshen Quotes site (libtv.md), Kimi was my first thought. Tight deadline, heavy workload—the rushed site won unanimous praise from fellow Mian disciples.
Anyway.
Other players already entered for the new season include the 2.4T parameter Qwen 3.8 Max. Sadly it launched just two days after K3—the LLM world's Wang Feng, forever overshadowed.
Nearing entry: GLM, DeepSeek, MiniMax, all rushing to train multi-trillion-parameter models.
The clearest progress is DeepSeek V4 Pro's official release. Though inconsistent—brilliant one moment, ghostly the next—it's 1.6T parameters. Given how Little Whale's Flash showed it could match post-training king GLM 5.2 with one-third the parameters, another small update and Little Whale could become the post-training king again, stunning everyone.
Once you grasp this season's basics, reaching conclusions is easy.
Kimi kicked off the new season with K3, leading for a month, followed by Qwen 3.8 entering the new season.
Meanwhile the old patch king Zhipu AI, relying on post-training sorcery to stubbornly extend its lifespan, dropped another GLM 5.3. That 753B base model took a full six months to optimize from barely usable to Fable alternative—not gonna lie, pretty magical, pretty hardworking.
Not to mention anything else, GLM 5.3 Macaron post-training edition is probably on the way. Looking forward to post-training of post-training of post-training.
I ran the ZangAI benchmark extra. GLM 5.3 did improve, but not by much—ranked behind Grok 4.6 and K3. Continues the good momentum, but lacks the groundbreaking impact of the previous version.
And maybe because it just raised prices and lifted Token plan purchase limits, GLM 5.3 runs tasks rather slow, averaging 12.7 minutes per task—far above DeepSeek V4 Pro (7.9 min) and Grok 4.6 (5.4 min).
By the way, DeepSeek V4 Pro official release: my runs show improvement over Flash, but not by much. The real big gainer is Grok 4.6—fast, high-scoring, and the full test run cost only 15.7 RMB, less than half GLM 5.3's API cost (38.2 RMB, calculated at 5.2 pricing).
Saint Ma's got manufacturing figured out. LLMs are pure manufacturing too.
That said, nothing against GLM 5.3, but this model kept teasing release without dropping, hyping ARR (we'll see the half-year report in a few days), plus the GLM 5.3 API leak last week—I'm genuinely curious if they can post-train their way to a GLM 5.4?
New season means: patch changed. Being #1 in the old patch is just extending your life. What decides new season rankings is shipping a 2T parameter model ASAP.
Rumors abound on timing. At minimum we can confirm MiniMax and GLM likely can't release next-gen models by month-end. Bear with it, fam—let Saint Yang enjoy his moment a while longer.
Kimi, temporarily in the lead, has two main domestic competitors.
One: Qwen will soon update a small patch. Qwen 3.8 Max's weakness is multimodal capability, plus the mindshare confusion from consecutive preview and official releases, preventing it from breaking through user mindshare.
Evidence: Qwen 3.8 Max's open-source version excludes multimodal capability, suggesting it may have Frankensteined on a separate VL small model.
The other: the mercurial DeepSeek. DeepSeek V4 Pro's actual performance deviates massively from its official benchmark results. Because official scores were run in DeepSeek Harness's minimal mode.
And DeepSeek Harness is explicitly a preview for developer testing, version 0.1, with overly novel design philosophy and interaction. So obviously few users actually use it. Thus DeepSeek V4 Pro official release, like its Harness, counts as half-baked.
The clout-chasing speed is genuinely fast
On the bright side: Qwen 3.8 Max and DeepSeek V4 Pro's next small patches could both potentially reproduce the post-training miracle.
Of course, LLM training is an engineering problem. No miracles, no secrets.
What I mean is, the post-training king title is like a model version number—valid for one month. Kimi and Qwen could both become post-training kings.
As domestic LLMs become pure manufacturing, abroad there's also the strongest opponent: the last saint of American manufacturing, Elon Musk.
Fable 5 and 5.6 Sol are rumored to be nearly 10T parameters—the sun and moon hanging over all models.
But flying too close to the sun means American LLM teams lack a hundred schools contending. The twin stars dominate, and even their relatively smaller models get slain by Qwen 3.8-27B—a model that runs on a Mac Mini, matching Opus 4.6 scores. You believe that?
Grok 4.6 is a 2T parameter model. My usage puts it at the same level as these top domestic models, and quite cost-effective—Grok API is cheaper than GLM, cheaper than K3. And Musk teased Grok 4.7 releasing in weeks: 2.1T parameters, officially entering the same season as K3 and Qwen 3.8 Max.
But seriously, dear Saint Ma loves our patch child Kimi a bit too much. I was scrolling Xiaohongshu and realized Grok bot is suspiciously Kimi-like.
If StepFun's logo can learn from Claude, then my Saint Ma can learn from Kimi. China and US learning from each other 👍
Saint Ma's journey also proves: LLMs are manufacturing without miracles. Saint Ma was showing off earlier, trying to replicate his Tesla/Twitter micromanagement-to-the-person style. Turns out it doesn't work—pure Hang Chen behavior.
Once Saint Ma stopped showing off, Grok got good. Logic is simple: do data honestly, train models diligently, seek truth from facts, no Great Leap Forward, no constant micromanagement and showing off.
Won't expand here—tomorrow I'll dedicate a piece to why LLMs are pure manufacturing.
One-sentence summary of Saint Ma's actions: distilling Chinese models, not distilling chain-of-thought, but distilling organizational approach.
That said: Saint Ma buying Cursor不如 buying Kimi. I see Kimi doing well—massively enriched features, basically eliminated model gap, productivity and entertainment both getting attention; if you added unlimited compute, Kimi would be our ideal Grok 👍
Then there's the literally evil capitalist Mark Zuckerberg stirring, whose new model from benchmark scores also shows fight, attempting to get the mouse back on the table.
In my previous piece on this miraculous company Kimi reads the sky for its meals. Back then K3 hadn't released, and K2.7 code contemporaneous with GLM 5.2 was an expensive, non-leading model.
I wrongly described this divinely blessed company: Kimi is the patch child of the model layer—neither creates patches nor falls behind them. When market attention shifts from the model layer, Kimi fades with the patch; when the model layer returns to hype, Kimi rises to greatness with the patch.
But K3 genuinely created a major patch. Looking back at these wild two months—from GLM 5.2 moonshotting Opus 4.8 to end the old season, to K3 forcing out Opus 5 and creating the new season.
The biggest winner is open-source models first. More specifically, current patch child is Kimi—larger base overwhelms. Meanwhile the other three baby dragons: Zhipu AI alongside MiniMax and StepFun all need to憋 something big at the 2T parameter level to prove themselves.
Finally, tying back to the title.
Ganlu Temple signifies Zhen Huan's bull market moment arriving. She endured hardship and suffering; her reward was true love with Prince Guo. Then true love shattered like a dream, and she resolved to return to the palace, confront Fat Orange, and reclaim power's center.
What I mean is, Kimi creating the new season this wave is the Dragon King returning, the bull turning, Zhen Huan back in the palace—the second round of palace intrigue begins, pure big-female-lead revenge drama.
(Cover image generated by ChatGPT, purely human-written.)
Welcome to our Bench hackathon this weekend to test K3 live, ample API supply. Scan the poster QR code to register ❤️