Come to the Internet Cafe and Put Every LLM to the Real-World Test
Everyone's an AI Model Whisperer Now
"Everyone's a Benchmark Emperor"
Recently, Zang AI, along with our friends Zhong Jingwei and Zhengyang, launched a product called Dongmodi (懂模帝). Anyone can now visit fuckai.md to browse all kinds of benchmarks.

We built this because we believe benchmarking is going mainstream.
In the past, when people evaluated models, they mainly looked at launch events and public leaderboards.
But now, one big comprehensive leaderboard isn't enough anymore. People care about whether a model actually fits their specific tasks, which scenarios it handles reliably, and where it falls apart.
Everyone has their own work and life contexts — and in their domain of expertise, they hold the power to define what capabilities matter.
Even Chinese people can fly ✈️
So having only AI application companies participate is far from enough. We want to find more people with benchmark-building skills and heavy Agent users to really put these foundation model companies through their paces.
So we thought of Zang AI's signature event: the internet café hackathon.
I hereby announce that the first Dongmodi Benchmark Hackathon will be held at a Beijing internet café. This event is co-hosted by Qwen and Zang AI.
Still 30 hours. Still unlimited instant noodles. Still no smoking.
The only regrettable thing is, since internet cafés ban minors, we might miss out on the AI world's prodigy teenagers 😭
This hackathon has two tracks:
The first is the Useful Track. You can take problems you encounter in daily life and work, and use them to test models.
Especially those scenarios where various companies "hyped things up massively," but the reality was complete garbage. We want the problems to be as hard for models as possible.
These could be life tasks like schedule planning, to-do organization, travel itinerary planning, apartment hunting, stock trading advice; or work tasks like dashboard building, data analysis, invoice sorting, and so on.
We especially welcome game development and video editing problems.
And of course, if you really love your job, you can focus on testing models' professional abilities in Coding (real bug fixes, legacy code archaeology and refactoring, preparing a mountain of shit code to see which model can handle it) and Cowork.
The second is the Meme Track. You can let your imagination run wild and build eye-catching, fun evaluation sets that probe models' underlying capabilities.
For example, having an Agent roleplay as a young drifter in Beijing trying to survive, playing Civilization or Dating Simulator, playing Werewolf or poker tournaments, or a Minecraft Raiden Shogun reconstruction competition.

Roleplaying as a young drifter in Beijing

Having an Agent play Civilization
The competition has two phases:
The first phase is building a benchmark. A benchmark consists of a set of problems, with explanations of why it's effective, what model capabilities it tests, and the scoring criteria.
For this competition, quality far outweighs quantity — you typically don't need more than 15 problems. With the right type and quality, even a single problem can score very highly.
This benchmark could be an environment where an Agent acts as CEO running a company, a tool for a specific game match, or automatically scored problems using programs and models.
The second phase is running the benchmark. Contestants need to run their evaluation sets on five models: Claude Opus5, GPT5.6 Sol, Qwen3.8 Max, K3, and DeepSeek V4 Pro, and submit model outputs, analysis reports, and scoring processes.
We'll provide detailed operation guides.
As you can imagine, both phases will consume massive amounts of tokens, and we're now in the era of major token price hikes.
Not even mentioning Claude and GPT's native prices, even DeepSeek's Saint Liang has turned into Prison Liang and started harvesting. Xianyu now spends two thousand RMB just to run one Zang AI Benchmark once.
Here we must thank Qwen and Moonshot AI for their sponsorship support.
For building benchmark scenarios, we recommend using Qoder and Qwen Office as your main harness, though you're free to use others;
For the benchmark-running phase, they're also providing each contestant with sufficient API credits to test the specified models.
If you have no idea how to build a benchmark but have a clear work scenario you'd like to evaluate, you're also very welcome to join our benchmark hackathon.
Evolvent AI is a professional benchmark-building data company. They'll be on-site to teach everyone how to build a good benchmark from scratch.
ZhenFund and we have prepared practical prizes for contestants including products from Bambu Lab, DJI, and Insta360, and have purchased a Mac mini as the grand prize.
Qwen has also added a special Qwen's Choice Award. If your benchmark surfaces capability blind spots that Qwen's internal evaluations haven't covered, you can win a DJI drone from Qwen, plus the opportunity to be officially included in Qwen's benchmark suite.
Quality benchmarks produced by contestants will also be featured in Dongmodi's curated recommendations afterward, with dedicated coverage and promotion.
Additionally, we've arranged learning sessions on-site, with定向邀请 AI application teams, model users, and benchmarking practitioners to share how they built their benchmarks, why they ultimately chose certain models, and which capabilities don't show up on public leaderboards but are critical in real business scenarios.
We especially welcome current students to participate. Outstanding performers may get "BOSS Zhipin'd" on the spot.
Event notes:
-
Internet café seats are limited and will be prioritized for contestants. Each contestant is also advised to bring their own laptop, unless your benchmark only runs domestic models.
-
To prevent another incident where hackers at the internet café stole contestants' accounts, every contestant must bring their ID card, and must log out before leaving the internet café 😭
If you're interested in our event, click "read more" or scan the QR code on the poster to register.

Time: August 22, 12:00 to August 23, 18:00
Location: Beijing (to be notified after registration)
Spots are first-come, first-served, until full. Substack channel registrations get some priority.
Thanks again to Qwen, Moonshot AI, ZhenFund, and Evolvent AI for their generous support of this event.
(This article's cover image was generated by ChatGPT; purely human-written content)
⬇️
Subscribe to our Substack: funeralai.substack.com