Doubao Lands Jobs with Multimodal Coding Skills
Multimodal Coding Upgrade
"Multimodal Coding Upgrade"
I wrote before that once Doubao Work launched with Seed 2.1 pro, it was already capable enough for office work.
But what are these internet giants best at? Creating demand where none exists.
Recently Seed dropped yet another update — a "minor" version called 0915. Except it wasn't just bumping parameters and calling it a day. It integrated multimodal capabilities with Coding & Agent features.
In plain terms: Doubao got better at coding tasks that require actually seeing, and at getting work done 👍
Right now the AI world is buzzing about three things — 3D generation, personal assistants, and long-horizon tasks — all of which depend heavily on multimodal input and output. So let's put the updated Seed-2.1-pro-0915 through its paces ⬇️
01
Doubao Does Interior Design
First, I tested Seed 2.1 pro's ability to understand and reconstruct the physical world. I took photos of my living room from three different angles and sent them to Seed 2.1 pro, asking it to build a model of my space.


Worth noting: these four photos were taken casually from random angles with no consistent pattern, so the AI had to figure out the orientation of walls and objects itself — not easy.
But Seed 2.1 pro pulled it off ⬇️

Furniture layout: correct. Rug distinguished from floor: check. It even recreated my bay window.
Then I asked Seed 2.1 pro to generate four different renovation schemes based on the original 3D model, with renderings and price quotes.
It first searched the web and gave me detailed furniture purchase plans ⬇️

Then it turned the whole renovation plan into an interactive dynamic webpage for a more intuitive preview ⬇️
Seed 2.1 pro's taste is genuinely on point. Combined with Seedream 5.0's image generation, renting in a city like Beijing where landlords are basically inhuman — just have AI whip up a scheme, spend a little money, and your quality of life skyrockets.
Honestly, if Mark Zuckerberg had had this trick back in the day, metaverse real estate would've pumped another 50 years 👍
02
GeoGuessr
The focus of this Seed 2.1 pro update is visual understanding, so GeoGuessr is the perfect test.
Quick explainer: GeoGuessr is a global street-view guessing game made by Swedish developer Anton Wallén in 2013. The rules drop you randomly somewhere in the world in 360-degree street view. You can pan around, move forward and back, look for clues, and deduce your location.
Kind of like the movie Searching — figure out where you are from images.
So if Seed 2.1 pro can crack GeoGuessr, it proves it actually has chops in visual understanding and information retrieval.
Let's be real, if you asked me to play GeoGuessr manually, I'd be guessing blind.
So I opened the Chinese version of GeoGuessr (Tuxun), sent Seed 2.1 pro a massive collection of strategy guides compiled by veteran players, and had it do deep learning and independent thinking. After a while it distilled its own Skills ⬇️

Then I just had it start matches on its own.
Unfortunately, in the first few rounds, Seed 2.1 pro either guessed wrong or took so long manipulating the webpage that it timed out and scored zero. I even got taunted by opponents a few times.
Fortunately Seed 2.1 pro has a spirit of criticism and self-criticism. After each round it summarized its mistakes and wrote new rules into its Skills.
After a dozen or so rounds, Seed 2.1 pro actually became a Tuxun expert ⬇️
Not only did it finish within the time limit, but its pin landed just 15 kilometers from the correct answer.
15 kilometers — I travel farther than that to meet friends in Beijing.
And from then on, it consistently kept its guesses within 200 kilometers.
I genuinely think this feature deserves to be spun out as its own app. Yong Ge's thing makes you do a 360 spin just to know if a location works for opening a shop. Seed takes one glance and knows your latitude and longitude.
Going forward, I recommend downloading Doubao before traveling to Southeast Asia. If you get kidnapped, snap a few photos from inside the compound — rescue success rate goes straight up 🤭
But to be honest, this test succeeding also owes a lot to Harness's Codex: the computer use is insanely fast and smooth, to the point where I suspect Sam Altman hired tens of thousands of African workers to operate it behind the scenes.
Maybe that's where the next AI battle will be fought.
03
Cloning a Douyin Mini-Game
Next test: if I give Seed 2.1 pro a game trailer, can it recreate the actual game?
Douyin is full of those scammy mini-game ads where what you see is nothing like what you get. I randomly grabbed one ⬇️
Then I sent Seed 2.1 pro exactly one sentence: "Based on this mini-game ad, make a vertical web game demo."
I didn't touch anything midway. Here's what Seed 2.1 pro built on its own ⬇️
Since it's pure frontend, the art is a bit rough. But this is exactly what I imagined the mini-game should look like.
I played it a few times and didn't find any bugs. Feels like it could go straight to market as a casual game and start making money.
Kind of terrifying when you think about it. No wonder GTA guards its trailers so closely — now AI can connect directly to Blender and Unity via MCP. Anyone with too many tokens at home could reconstruct GTA 6 from a trailer
Can't help but marvel: just a year ago, people were still founding Vibe Gaming startups, treating pixelated side-scrollers like precious gems. Now you chat with a model for two minutes and anyone can be Shigeru Miyamoto or Hideo Kojima. Who do you even complain to?
All I can say is: it's fast. The g-force is real 😰
04
Making a Trailer for the Mini-Game
Finally, let's reverse the test: feed the mini-game from the last test back into Seed 2.1 pro and have it create a trailer suitable for ad placement.
Of course, this isn't just writing a prompt and having Seedance render it — that would be too easy on Doubao.
I demanded that Seed 2.1 pro write a script with maximally complex camera work, then build white-box models in Blender for each scene, produce a complete previs video with camera moves, feed that white-box previs as motion reference into Seedance 2.5, have it lock camera trajectories and event timing, repaint the white boxes into Hollywood blockbuster-style footage, use the first shot as a style anchor to ensure visual consistency across the whole piece, and finally edit it into a roughly 379-second vertical trailer.
I must stress: generating white-box video before rendering isn't pure showboating. With complex camera work and multiple motions, editing white boxes is genuinely more precise and polished than generating directly from prompts. That's why every video model is pushing this feature now. I used to be skeptical; after trying it, I get it.
Anyway, here's the script and white-box video that Seed 2.1 pro generated on its own ⬇️
To be honest, I couldn't really follow this video. Especially that thunder cylinder at the beginning — no idea what it's supposed to mean.
But then I thought: this isn't for me anyway. AI makes it, AI watches it, and I'm just there to click approve. I felt at peace.
Sure enough, once Seedance 2.5 rendered the video, everything came together ⬇️
Back in the day, generating something like this would've taken maybe five rerolls per shot. And Seedance 2.5 is expensive — generating blindly like a headless fly means Doubao laughs all the way to the bank while you go bankrupt.
But Seedance 2.5 is expensive; Seed 2.1 pro is cheap.
So from now on, use Seed 2.1 pro for white-box previs first — that's indirectly driving Seedance 2.5's price down 🤫
After all these tests, I think the real question Seed 2.1 pro's update and multimodal capabilities are solving is: to get AI to do work, how much work do we have to do for it first?
Right? If it's like before — where using AI meant organizing materials, describing scenes, then checking and redoing — then AI is just like a robot vacuum: no more sweeping, but you're still cleaning the water tank and brushes.
Labor doesn't disappear, it just shifts. Users do pure busywork, pure waste of electricity and money 😭
So now everyone's racing on multimodal. Fable 5.1 emphasizes visual checking. GPT-6 Astra supports computer use. Even DeepSeek patched multimodal input into V4. If LLMs want to do jobs start to finish, multimodal is unavoidable.
Beyond saving users effort, multimodal capability itself strongly correlates with output accuracy. As they say, button the first button of life right: if there's information loss at step one, how can you expect big results from the model?
So with this Seed-2.1-pro-0915 update, stronger multimodal means Coding and Agent capabilities leveled up too. Doubao is really going for that hexagonal warrior status now 👍
(Cover image generated by ChatGPT; article written entirely by human)
⬇️
Subscribe to our Substack: funeralai.substack.com