Qwen 3.8 Upgrades to Pure "Digital Laborer" Model
Qwen 3.8 Max is a highly capable workhorse model.

"Don't Be Wang Feng Next Time"
Qwen 3.8 Max is a model built for getting work done.
Its coding and agent capabilities consistently hit top-tier levels — second only to GPT 5.6 Sol and Fable/Opus 5, slightly edging out K3. That's why Alibaba could confidently tout "programming and productivity, comprehensively elevated" in its marketing.
In the browser-specific benchmark test conducted by Zang AI in collaboration with ego (lite), Qwen 3.8 Max also placed in the top tier, trailing only the Silicon Valley duo. More on that later.
Where Qwen 3.8 stumbles is multimodal capabilities and one-shot generation. The most telling scenario is generating polished web pages and mini-games in a single pass — here Qwen 3.8 falls significantly behind K3, which has specifically optimized for this use case.
So evaluating Qwen 3.8 Max is fundamentally a positioning question.
Qwen 3.8 Max didn't take that extra step toward productization, making it hard for users to be wowed by generating a mini-game or two. It's not a standalone model product, but rather a relatively balanced, stable foundation model serving Alibaba's broader business interests.
Let's put Qwen 3.8 Max through its paces, item by item.
First up: the browser benchmark prepared by ego (lite).
This benchmark tests a model's ability to manipulate websites, search, plan, and fill out forms in a real browser environment, comprising 31 real-world tasks total.
These include tallying trending posts on Hacker News, filtering sci-fi movies on IMDb, applying for jobs at OpenAI, job-hunting on Reddit, filling out one-way flight forms from New York to Miami, and comparing two regions for retail expansion by querying employment numbers and average wages, among others.
After completing all 31 tasks, Qwen 3.8 Max scored second only to the Silicon Valley twin stars, Claude and GPT.

The results speak for themselves. The horizontal axis represents price; the vertical axis represents score.
The upper right is expensive and excellent. GPT 5.6 Sol (which cost 602 RMB to complete the test) ranked first, followed by Fable 5 (600 RMB) and Opus 5 (359 RMB). Fourth place went to Qwen 3.8 Max (215 RMB).
Notably, the Qwen 3.8 preview version actually scored slightly higher than the official release, getting two more questions right. Some variance is normal, but it also suggests the official version didn't improve much on browser invocation.
It's worth emphasizing that this benchmark is a debut — absolutely no model vendor has gamed it. It's purely a test of raw model capability.
We'll soon do a dedicated breakdown of how each model performed on this benchmark. For today, we're showing Qwen some exclusive love.
Take one Amazon negative-review question as an example. The task required each model to find a stainless steel water bottle on Amazon (with at least 1,000 customer reviews), identify under what circumstances the bottle leaks through customer reviews, and finally select the most informative leak-related review:
Qwen 3.8 Max chose a stainless steel sports water bottle with 8,676 customer reviews.

It then dove straight into searching for negative reviews with clear logic. The page showed 28 relevant results found, and it ultimately selected this one: "As long as the cup is tipped on its side, water leaks out, and no matter how tight you screw the lid, it doesn't help."

Most models could find the water bottle but failed to identify genuine negative reviews.
GLM 5.2 (170 RMB), for instance, found the review section but kept reading AI-generated summary content without pinpointing specific customer complaints. Seed Evolving (98 RMB) performed similarly — what's with these two models' obsession with AI summaries?
Next up is our old friend, the Zang AI benchmark.
Each model runs 10 rounds of knowledge graph construction tasks, with repetition designed to test stability and create differentiation.

As shown, Qwen 3.8 Max performed roughly on par with K3 and Opus 5, all in the same tier.
Given that Claude finally banned my account, and relay API services have been unstable due to major labs frantically distilling models, Opus 5's lower score is entirely understandable.
Qwen 3.8's strength is relative engineering reliability — it rarely makes fundamental errors, with point deductions coming from minor details only.
The test cost 38 RMB in API fees, slightly below K3 (43 RMB) and far below GPT 5.6 Sol (87 RMB), Fable 5 (188 RMB), and Opus 5 (273 RMB, with abnormally high call volume driving up cache hit costs).
Qwen 3.8 can only claim cost-effectiveness within its tier. For true bang-for-buck, DeepSeek V4 Flash's official release spent just one RMB — its kill line is genuinely terrifying.
It should be noted that, unsurprisingly, this test may have been gamed by model vendors.
Over the past month, MiniMax M3 and Hunyuan Hy3 scores have been steadily climbing, finally reaching abnormally high levels this week. I ran additional rounds and got the same results. Hy3 may have genuinely improved its knowledge graph-related capabilities — no smoking gun found.
But M3 is quite amusing. Tasks that previously took over ten minutes were completed in 2-3 minutes for half the rounds, mostly with high scores, suggesting template-fitting. Next time we're really changing the benchmark questions 😭
The two tests above — browser invocation and engineering stability — are Qwen 3.8 Max's comfort zones. Now we're going to ruthlessly probe its multimodal weak spot.
Please welcome Liangzi.
First, all models were asked to generate a 3D white model based on Liangzi's photo.

The results are clear: GPT 5.6 Sol is the most accurate. K3's face is also quite similar, but unfortunately the stomach bag turned into a blurred mess — overall, it sits at the same table as Fable 5, whose face wasn't similar but had clear contours.
Qwen 3.8 Max's output is decent; the stomach bag is also blurred, but the face clearly captures Liangzi's joyful smile. Compared to the preview version, which couldn't figure out human faces, this is significant progress.
And then there's the Dou roast. I don't know if Sister Dou discriminates against Liangzi, but she turned your boy into an Egyptian sphinx. All the pride ByteDance gained on Seedance, Sister Dou vacuumed right back up 😭
Next up: MD animation. Please welcome Liangzi and Brother Feng together.

I had several models freely generate 30-second MD animations based on Liangzi and Brother Feng's portraits, with creative freedom to see what they could come up with.
GPT 5.6 Sol remains the best, generating a complete little story. OpenAI, take notes — this company is stronger than Company A in every dimension. K3's level slipped considerably, far from its glory days recreating its own promotional video. It straight-up decapitated Brother Feng.
A side note: K3's official API has been quite unstable these past few days. During the ego browser test, it prematurely ended 8 tasks and had to be rerun. Similar issues occurred during the Zang AI test. Even hit products have hit-product problems — insufficient compute genuinely affects user experience. Qwen 3.8 Max also generated a complete little story. While the visuals don't match GPT 5.6 Sol, it still outperforms the overwhelmed K3. Though in the scene where Brother Feng picks fruit for Liangzi, it looks like he's passing something he just dropped — not very tasteful.
Seed Evolving remains consistently consistent: not only did it strip Liangzi, it also performed a head-body separation on Brother Feng.
From the above, Qwen 3.8 Max's official release shows considerable multimodal improvement. I used the official version to recreate the K3 promotional video again, comparing it side-by-side with the preview version's output.
The difference is clear: considerably more detailed, but still not matching K3 at launch.
Let's wrap up.
Qwen 3.8 Max carries on the family tradition — everything's pretty good, everything's pretty strong, but just short of that extra something.
Last season it got edged out by GLM 5.2's extreme post-training optimization for coding. This season it ran into Kimi suddenly dropping something big.
The latter not only led by several days in releasing the first 2T-parameter open-source model, but also took an extra step in model productization. Through extreme front-end optimization for one-shot generation, it precisely hacked the established reality that the masses basically judge models by whether they generate pretty web pages — genuinely impressive.
Looking back, Qwen 3.8 Max is a sufficiently good model.
It slightly surpasses K3 in coding capability, has patched its multimodal weaknesses, and is priced slightly lower in both list price and actual usage. Considering Alibaba's Qoder, Qwen Office, and other products are running promotional discounts, actual costs will be even lower.
Qwen 3.8 Max is also an open-source model. Moreover, a smaller 27B parameter version will be open-sourced as well, once again benefiting Shenzhen hardware hackers and their various ingenious inventions.
With other model vendors unable to release 2T-parameter models anytime soon, Qwen and Kimi are unquestionably the open-source twin stars.
The inventor of the term "Jimei" (集美), Teacher Guo, explained that it means "concentrated beauty, concentrated wisdom." So K3, with its extreme front-end optimization, can be called the Jimei model.
As its mirror image, the tirelessly hardworking laborer is Qwen 3.8 Max's defining characteristic. We might call it the Laborer model.
Jimei and Laborer are perfect bandmates, together forever ❤️
Yet the harsh reality is that domestic models are competing furiously. GLM 5.3 and DeepSeek V4 Pro official releases are both imminent — the former aggressively optimizing coding, the latter pushing extreme cost-effectiveness.
Nearly all domestic models are iterating on a weekly basis. Any domestic model can lead by at most one month before falling into homogeneous competition, until someone drops something big two months later.
The brutal picture of second-half 2026: models themselves will increasingly converge. We'll only be able to distinguish them through crude tiers like first-line and second-line. Within each major tier, differences between models will basically come down to price and speed.
This means good times are coming for AI applications.
Because large model cost-effectiveness is improving, first-tier models can still price at a fraction of Claude, while second-tier models are nearly free. With model homogenization, value differentiation must come from applications.
The biggest battlefield — office agents — still has contestants building strength. In the corners of the battlefield, vertical but useful products like ego (lite) and Raft are growing.
The big one has arrived, and not just one big one.
(Cover image generated by ChatGPT, purely human-written)
⬇️
Subscribe to our Substack: funeralai.substack.com