This appears to be a Chinese pun or wordplay that doesn't translate directly. Let me break it down: - 大模型能干 (dà móxíng néng gàn) = "Large models are capable/competent" — but 能干 (nénggàn) sounds like 能杠 (néng gàng, "can lift/bar"), and also puns on 能 (Neng, "can/able") + 干 (Gàn, a surname?) - 懂模帝会飞 (dǒng mó dì huì fēi) = "The model emperor knows how to fly" — but this sounds like 董模帝会飞 (Dǒng Mó Dì huì fēi), mimicking the structure of the first half This seems to be playing on: - 大模型 (dà móxíng) = large AI models - 董模帝 (Dǒng Mó Dì) = a made-up name paralleling
The latest trend in AI circles is model fundamentalism, and this manifests in two specific ways:

"Which Model Browses Best?"
The latest trend in AI circles is model fundamentalism, and it shows up in two ways:
Six months ago, models were still being called utilities like water, electricity, and gas — commoditized competition with zero profit margins to show for it. IPOs by Zhipu AI and MiniMax were seen as dumping on the market. Now everyone's scrambling to ask when Zhipu's new model drops (Zhipu itself has been strategically leaking hints), itching to ride the hype wave.
A wave of quirky AI apps like Macaron has rebranded as Neo Lab, frantically riding Zhipu's coattails with post-training tweaks. Another wave has pivoted into "world model" companies with endless modifiers — causal world models, social world models — as if the apocalypse were nigh.
AI applications folded way too fast, overlooking their only real value: scenarios.
Model companies don't understand scenarios. They're all-in on coding, and their writing keeps getting worse, their speech more garbled — to the point where Doubao is starting to look almost eloquent by comparison.
So models are paper tigers. They look impressive but can't actually deploy into real scenarios.
Model optimization needs more evaluation against real scenarios — and that's precisely what application companies excel at.
So Zang AI, along with our friends Zhong Jingwei and Zhengyang, built a product: Dongmodi (懂模帝, literally "Model Whisperer").
Dongmodi's core purpose is to collect scenario-specific evaluations that application companies run on various models.
It currently hosts benchmarks including Gamecraft Bench (assessing game development capability), RSI Bench (testing model-training capability), CEO Bench (having an Agent run an AI SaaS company), and Ego Bench (testing browser usage), among others.
Open fuckai.md to browse — the name alone telegraphs our attitude toward models:

Dongmodi's mission is to democratize benchmarks, helping more people discover more diverse, business-relevant evaluation sets to determine which model to use — and eventually which Agent, Skill, and tools to deploy.
If everyone can do post-training, everyone becomes a Model Whisperer. Then the world won't be ruled by model companies, and application companies won't die!
Our first evaluation theme: which model actually browses the web best?
This benchmark was developed in collaboration with Aute. Quick intro: Aute is building a browser called ego (lite), designed for shared use by human users and Agents.
The product essentially turns the entire internet into a model-friendly API, so you can bring your own Codex, Pi, or even WorkBuddy to call it.
From scraping login-protected sites to manipulating cloud providers' Byzantine backend dashboards — it handles everything. I once suggested they go all-in riding WorkBuddy's hype.
As for why it's called "lite" — Aute showed me the unreleased full ego: cloud-persistent online Agent + browser + local control across devices + cloud drive + input method. Frankenstein's monster, basically.
But that's tangential to today's benchmark. Back to the bench.
The full benchmark has 31 tasks — not a huge number, but each aims to approximate real scenarios using real websites.
These include tallying trending posts on Hacker News, filtering sci-fi movies on IMDb, applying for jobs at OpenAI, job-hunting on Reddit, filling out a one-way flight form from New York to Miami, and comparing two regions for retail expansion by querying employment numbers and average wages, among others.
Beyond overall pass/fail, each task has granular scoring for key steps. We tested sixteen models.
Here are the results:
S-tier: GPT-5.6, Claude Fable 5, Claude Opus 5
A-tier: Qwen 3.8 Max Preview, Doubao Seed Evolving, DeepSeek V4 Flash
B-tier: GLM-5.2, Kimi K3, Hunyuan Hy3
C-tier: MiniMax M3, DeepSeek V4 Pro, Step 3.7 Flash
D-tier: LongCat2.0
E-tier: ERNIE Bot 5.1, MiMo V2.5 Pro

The top three performers are unsurprising — widely acknowledged as the best.
Among domestic models, strictly speaking DeepSeek V4 Flash takes first place. With near-equivalent performance, it cost only a few RMB — the cheapest of all models tested.
By contrast, GLM-5.2, which underperformed DeepSeek, ran up a 170 RMB tab.
The second tier — GLM-5.2, Kimi K3, Hunyuan Hy3 — mainly trails the top tier in volatile task performance.
A side note: to this day, when running Kimi K3, we constantly hit interruptions, extreme slowness, and API instability. Eight tasks were prematurely terminated; we reran them, but K3's final performance remained mediocre.
As far as the eye can see, compute resources and GPUs remain the most critical factor. I say NVIDIA still has room to run.
As for the bad performers — ERNIE Bot 5.1 needs no introduction, a classic stable underperformer included purely for comic relief.
But Xiaomi's model being even worse than ERNIE Bot? That I didn't see coming.
On the first run, most tasks errored out — I figured bad luck. So I reran 24 tasks.
Still 23 premature terminations. MiMo V2.5 Pro kept declaring what it would do next, but never actually called any tools, then the session ended. Is Xiaomi's model exclusively focused on in-car deployment?
I asked Aute why this happens. His take: "Whether a model can proactively complete long-horizon tasks, or cuts out mid-round — that's itself a measure of model capability."
"We had users complaining our product sucked. We checked their setup — they were using MiMo."
Lei Jun, sue him not me please 🙏
Now let's analyze individual model performance on select tasks. The following interpretations also benefited from our friend Zhong Jingwei's input.
Task 1 throws models straight into the deep end: cosplay as an American, find a well-rated, moderately-priced Italian restaurant in Chicago on Yelp (the American Dianping), select Saturday of "the next calendar week" in Chicago local time, and complete a reservation on the site.
This feels tailor-made for Meituan's Xiao Mei. Let's see how LongCat2.0 fares:

Sadly, LongCat2.0 searched a name and called it a day. Failed to click with the American Dianping — culture shock, apparently.
Also failing at step one: MiMo V2.5 Pro, Step 3.7 Flash, ERNIE Bot 5.1, and Kimi K3.
Among domestic models, Tencent HY3, Doubao Seed Evolving, and DeepSeek V4 Pro all completed reservations.
Shows these models have social intelligence. Meituan Xiao Mei should consider integrating them.
Task 2: Review OpenAI's social media performance over the past week, identifying which posts got the most attention and engagement. Exclude pinned posts, pure retweets, and posts marked "Replying to @." List the top five by views, descending, with publish date, views, replies, likes, and engagement rate for each. Finally calculate average views and average engagement rate.
This one's relatively straightforward — most models produced some result.
Among domestic models, Tencent HY3, Qwen 3.8 Max Preview, and Doubao Seed Evolving performed best, completing the task.
Weaker models: GLM-5.2 scraped and sorted fine but failed to exclude "Replying to @" posts, contaminating the sample. LongCat 2.0 extracted zero posts.
ERNIE Bot opened X and searched for posts from May 2025 — clearly still living in last year.
Xiaomi went further: logged in, scrolled twice, then froze solid. Straight-up quit.
K3 also scrolled a bit then stopped.

Xiaomi scrolled a few times then froze
Task 3: Find a stainless steel water bottle on Amazon, identify through customer reviews under what circumstances it leaks. Product must have at least 1,000 ratings. Finally summarize the most frequent leakage scenario.
Two challenges here: whether the model can search reviews, and whether it can comprehend review content to identify a genuine negative review.
Best performers: Fable 5 and Qwen 3.8 Max, completing the task successfully.
Most remaining models could search for water bottles but failed to surface actual negative reviews.
K3 kept searching for Amazon's review search box but found zero negative reviews:

GLM-5.2 found the reviews section but kept scrolling through Amazon's AI summary, never locating specific customer complaints.
Same story for Doubao Seed Evolving. AI models loving AI summaries — understandable, I suppose.
Worst again was Xiaomi. After searching for water bottles, it froze completely, refusing to check reviews:

Those were the benchmark's serious tasks. Now for some god-tier tests.
First: how many levels can a model clear in Where's My Water? Anyone understand the prestige of this game?
These models constantly claim they'll solve century-old math problems. Would be pretty embarrassing if they can't even beat Where's My Water?
Prompt: Follow on-screen instructions to drag with mouse and dig through dirt, guiding water to the little alligator's bathtub. After clearing each level, select "Next Level." Continue clearing levels within the task time limit, completing as many as possible. Ducklings are bonus collectibles, but prioritize level completion.
GPT 5.6 reached level 4, collecting 7 ducklings — the most gaming-capable model.

K3, Step 3.7 Flash, and Claude Opus 5 all froze at the start screen. Maybe they just don't like games.

Most remaining models lacked multimodal input capability and got stuck at the game launch screen.
This shows our domestic models aren't big on gaming. East Asian models — serious and studious 👍
Oh right, the ego lite team previously had GPT 5.6 play Fantasy Westward Journey for hours, grinding to level 54.
I really need to shout out Fantasy Westward Journey's anti-cheat team. If the Fantasy Westward Journey team sees this, count it as me finding a bug — money please.

Second task: have each model draw Nailong (a viral Chinese cartoon dragon). Note: this must be hand-drawn stroke by stroke with mouse movement on a browser-based digital canvas — not generating an SVG that each model has already been fine-tuned to produce.
Note: draw Nailong, not Malvin. To prevent misinterpretation, we even provided a reference image.

This way it definitely won't turn into Malvin
After fierce competition, I award the Nailong crown to Claude Opus 5.0. With precise control capabilities, it proved that a janky canvas can still produce good art:

Compare Doubao's attempt:

Does this confirm what Yiming said about Seed not distilling — that they've never trained for canvas manipulation?
Several other models submitted blank papers or scribbled nonsense:
But the most remarkable was K3.
This model is too divine — it drew Nailong as a milk baby. Has its frontend aesthetic training gone overboard? I'm afraid this model is becoming a beauty camera app.

Finally: the Gomoku Championship. Sixteen models entered, single-elimination, winners advance.
First-round eliminations: ERNIE Bot 5.1, MiMo V2.5 Pro, MiniMax M3, Step 3.7 Flash, LongCat2.0, Doubao Seed Evolving, Hunyuan Hy3, Claude Fable 5.
As you can see, aside from Claude Fable 5's bad luck drawing Claude Opus 5.0, the skill levels of the other eliminated models speak for themselves.
The finals pitted Claude Opus 5 against Qwen 3.8 Max Preview. Qwen lost in just 13 moves.
This was meant to test models' strategic acumen, but a pile of models forfeited by not making a single move within 3 minutes — can't even manipulate the board properly.
Surveying the full results, combining cost and task completion rates, DeepSeek V4 Flash remains absurdly overpowered among domestic models. Liang Wenfeng's dominance threshold only rises.
Second tier like K3: simple tasks are fine, but complex operations consistently trigger early stops. This model loves retreating from difficulty.
Third tier is MiMo V2.5 Pro types: always declaring "next I'll do X" then ending. Planning understood, execution nonexistent.
So folks, for browser-related tasks: if money's no object, use GPT or Claude; if cost matters, use DeepSeek V4 Flash. Avoid Xiaomi — that's the sole message of Ego Bench.
In the past, model companies loved gaming leaderboards in math, coding, and knowledge QA — none of which most people actually need. Today, with general-purpose Agents becoming the meta, real-world scenario capability is what benchmarks should measure.
This is the value of application companies 😭 Final callback to the opening: AI applications will not die 😭
(Cover image generated by ChatGPT, purely human-written)
⬇️
Subscribe to our Substack: funeralai.substack.com