A Hackathon Killed the Whole Internet Café

Fermented mung bean milk tastes great.

"Fermented Mung Bean Milk Tastes Good"

Last weekend, we held our very corporate, very Beijing-flavored Dongmodi Bench Hackathon at an internet café, going hard on our mantra that "everyone should have the right to define what each model can do."

After a full day of intense coding, we produced over 30 benchmarks. The demo presentations alone lasted five hours — by the end, we were all getting a little oxygen-deprived.

Peak intellectual density at an internet café, truly.

Qwen generously provided every participant with RMB 1,500 worth of tokens to build and run their benchmarks. By the end of day one, only four people asked for more.

Turns out Qwen3.8-Max is built to last. Good stuff — use more of it.

Enough preamble. For those who missed it, here's the recap:

01

Qwen Loves You Back

Maybe Qwen's reputation as an open-source model preceded it. A contestant named Dongyun got a wild hair on the first night and straight-up deployed Qwen3.5-7B on an internet café PC.

When I asked him why, he said there was no particular reason — he just can't stand seeing idle hardware. He usually keeps a bunch of 30B models running on his own server just to pit them against each other.

He also said, "The GPUs here (RTX 4080s) are criminally underrated. Using them only for gaming is a waste. Running LLMs on them would be so much better."

The internet café PC met its strictest father.

One participant brought an embedded low-level driver development benchmark — the name was so technical I didn't understand it at first.

Turns out the setup gives the model a chip peripheral manual and a virtual dev board, then has it write drivers, read/write registers, compile and run, and handle interrupts/timeouts/exceptions. Both of his test cases were drawn from his actual work. Qwen3.8-Max ranked first on both, demonstrating remarkable system-level agent capabilities: reading manuals, reverse-engineering peripherals, building tests, diagnosing platform issues, and iteratively fixing bugs until closure.

The Qwen Choice Award ultimately went to Vivian from Meituan's Longmao team, whose day job is literally model evaluation.

After our event, she clocked 7.5 workdays this week — meaning, hey boss, overtime pay please.

Her benchmark suite was called "Everything Can Be Evaluated," encouraging people to draw test cases from their own work. Wait, isn't that our event slogan? Nesting dolls much?

Vivian's test cases covered: DeepSeek harness multi-plugin compatibility, whether a certain top podcast host plagiarized, distilled, or just localized content, Markdown multi-view evaluation, AI social research, and offline cataloging of Xiaohongshu image assets — the breadth of interests was something else.

02

The Food Delivery King Takes Home a Mac mini

On to the winners.

The grand prize Mac mini went to RainFly, who also won a PS5 at our last internet café hackathon. Dude's been restocking his inventory through us.

RainFly went deep into the masses with his "Food Delivery King" benchmark for comparing food delivery prices, mainly testing models' coupon-clipping abilities.

He mocked a Meituan App environment where the model had to find the optimal strategy across six entry points: daily coupons, free "god coupons," coupon inflation, daily red packets, store red packets, and pop-up coupons.

Per his evaluation, Qwen performed like competitive eater Liangzi — maxing out all three free benefits across three rounds, with decent prices too, though it took the longest time.

The second winner was Hao, who built a 911 dispatcher benchmark.

As the name suggests, it tasks models with dispatching public emergency units like ambulances and fire trucks under constrained resources.

Hao pulled from a public San Francisco dataset, extracting a 20-hour window where the LLM had to make decisions without seeing the future. Response time, service coverage, safety constraints, and resource efficiency were all auto-calculated.

This benchmark is almost too useful. I'm calling it the US Kill Line Bench.

The final winner was Zhang Zhenyao from robotics, who built a benchmark testing models' physics capabilities with problems including drone patrol routes, robotic arm grasping, and friction-based motion.

The inspiration came when he asked a model to write drone control code before the competition. The model spat something out fast. He dropped it into a simulator — crashed immediately on takeoff. Suspected Kobe aircraft code leak.

I propose we call this the World Model Bench.

Going forward, it can test all those world models that love boasting about mastering physics laws but are actually just Wan post-trained.

NemoVideo, an AI editing agent company, contributed a video editing benchmark. They loaded five models into the same editing agent, ran them ten times each on identical footage and prompts, and checked if they could independently cut a finished piece.

The NemoVideo folks told me they discovered some agent flaws while batch-running evaluations and started emergency fixes and iterations. The benchmark results are now being used by their engineering team as a reference for product roadmap planning.

This is exactly why we do benchmark content — to revive AI application companies!

03

Who Had the Most Unhinged Benchmark?

The shitpost track brought out all kinds of deities: a Dating bench, a Beijing Life bench, a career-change bench.

Looking forward to someone doing a Fat Cat bench next time.

Among the two winners, one built an entertainment operator benchmark.

Five models got identical budgets to manage three trainees, each with their own weaknesses, over eight weeks — training, staging performances, earning resources, handling emergencies, all to successfully debut them.

Kinda like the CEO bench we wrote about before.

The other winner built the "Bull Market Incoming" bench — slightly clickbaity name.

Because what it actually does is give different models the same "Bull Market Incoming" character persona and have them produce a video.

Unlike direct video generation, the model first uses code to create a runnable 3D world, establishing spatial and temporal rules for characters, scenes, lighting, and objects, then directs performances, schedules cameras, and completes filming within that world.

Upon learning the Bull Market bench won, one passionate participant who pulled an all-nighter and even brought an RTX PRO 6000 to run benchmarks lamented, "Bull gets the prize, missed the Bull Market dividend."

Proves that personal effort matters, but mainly it's about being chosen by the era.

04

Long-Horizon Tasks: Kimi

Kimi, another domestic shining star, also provided every participant with RMB 1,000 worth of tokens.

Let's see how K3 performed:

Take the aforementioned Bull Market bench, which tests the widest range of model capabilities. All models had to complete storyboarding, Three.js world building, character animation, sound design, browser debugging, and final delivery through code.

Kimi ranked first across six dimensions: camera work, character continuity, visual consistency, animation continuity, and originality. Hard to believe a language model could reach this level.

Please enjoy Kimi's Bull Market Incoming:

My friend @Digital Sheep contributed a fortune-telling benchmark. He selected 199 auto-gradable multiple-choice questions from a global fortune-teller competition and had five major models answer them.

Not about lottery prediction, of course — mainly testing their ability to call fortune-telling tools and make temporal judgments.

K3 made 344 tool calls, showing solid structured long tool chains. Final score was second only to Opus 5.0. The only downside: completion time was 51 minutes, more than double Opus 5.0's.

Shows K3 has clear advantages in long-horizon tasks that "follow existing paths."

Bigger is better, apparently. If compute allows, K3 probably keeps on dominating.

05

The Fermented Mung Bean Milk Immortal

Learning from our last vinegar fish experience, we bought dozens of bottles of fermented mung bean milk (douzhi) this time, convinced someone would finish them. We vastly underestimated douzhi's power:

Xianyu, after drinking, fumed that it was fermented fecal water. One participant, after finishing, was shocked to find the private room's smell wouldn't dissipate and had to leave.

Of course there were masters. The same guy who locally deployed Qwen chugged half a bottle without changing expression.

I have reasonable suspicion he's Beijing-born.

06

Contestant Becomes Net Café Admin

When I arrived on day two, I saw an extra person at the front desk. Figuring they'd called in temporary help, I asked the guy to put up two posters for me. He didn't refuse.

Later he set up a computer at the front desk and started coding. For a moment I thought the admin had joined the competition. Turns out he was always a contestant — the net café admin had temporarily drafted him to work.

I later asked what he was thinking. He said, "I pulled an all-nighter and had nothing to do, so I chatted with the morning-shift admin and helped out, cleaned up, set up the venue."

Next time I'm posting a "No Admin Enslavement of Contestants" sign.

07

The Matrix

This was the most hacker-dense internet café hackathon in our history.

During the final presentations, one participant accidentally hacked all the internet café computers while testing model safety capabilities — nobody could boot up 😭

One hackathon took down the entire net café.

We're now in a cold war with the net café owner. No more using this place, sad 😭

To wrap up: the reason we chose Claude Opus 5.0, GPT5.6 Sol, Qwen3.8-Max, K3, and DeepSeek V4 Pro — five acknowledged top-tier models — is that good benchmarks should create differentiation among top performers.

That said, overseas models underperformed this time. Unclear whether it was API instability or stealth dumbing-down during peak hours.

If everyone were testing second-tier models like Xiaomi MiMo or Meituan Longmao, we'd see clear winners and losers, defeating the purpose.

Though we can't rule out hosting an event specifically to test how bad these models are.

We'll also publish a dedicated analysis of the above benchmarks — stay tuned. Thanks to Qwen, Moonshot AI, ZhenFund, and Evolvent AI for their generous support.

(Cover image generated by ChatGPT, article 100% human-written)

⬇️

Subscribe to our Substack: funeralai.substack.com