"Make Agents Easy for Everyone" | A Conversation with Zhang Xiantao, President of Alibaba Cloud Wuying
Anticipate shifts in advance and prepare for the changes coming over the next two to three years.
Anticipate changes in Chinese ahead of time, and prepare for the shifts two to three years down the road.

👦🏻 Podcast interview: Koji, Ronghui
🥷 Edited by: Starry
🧑🎨 Layout: NCon

China's big tech executives are a hidden talent pool — formidable but generally low-profile, rarely stepping out for long-form podcast interviews.
This week, Crossing welcomes precisely such a heavyweight who has spent years working behind the scenes: Zhang Xiantao (alias: Xuqing), President of Alibaba Cloud's Wuying (Shadowless) Business Unit.

He previously led Alibaba Cloud's Elastic Computing team for years, guiding them through the design and implementation of the "X-Dragon" architecture. This architecture established Alibaba Cloud's distinctive position in the global cloud computing industry and earned broad international recognition at technical conferences and in academic papers.
Xuqing's topic with us: Agent Infra.

2025 is the Year of the Agent — eight out of ten entrepreneurs are building one. Crossing has conducted extensive interviews and evaluations on this over the past few months, and found that what determines an Agent's capability ceiling isn't just the model itself, nor just engineering and interaction polish. Infrastructure matters too — Agent Infra.
Memory, Tool Use, Task Planning, Runtime, Multi-Agent collaboration, security and privacy mechanisms — every piece is indispensable. It's like when a new employee joins: you need to give them a computer, network access, and tools like DingTalk, Lark, email, and various AI productivity tools. Without proper infrastructure, both Agents and new hires end up just sitting there staring blankly.
Where there are challenges, there are opportunities — which is why both startups and big tech are pouring resources into Agent Infra. Alibaba Cloud's Wuying team launched AgentBay as an entirely new experiment: providing Agents with cloud sandbox environments, compute power, and toolchains to build a complete runtime. Behind a recently viral global-first universal mobile Agent, Wuying provided the Agent Infra.
This week, we welcome:
Zhang Xiantao (Xuqing) | President, Alibaba Cloud Wuying Business Unit
Qu Liwei (Anchen) | Product Lead, Alibaba Cloud AgentBay
Together, we'll explore the past, present, and future of Agent Infra — from technical principles to industry landscape, from strategic judgment to organizational transformation, to maps of entrepreneurial and investment opportunities.
If you're a developer, founder, or investor, this is the episode that will help you truly understand Agent Infra.
Listen on WeChat:
Listen on Xiaoyuzhou:


👦🏻 Koji
2025 is the Year of the Agent. Eight out of ten entrepreneurs are building Agents of one kind or another. At Crossing, we've done extensive interviews and evaluations over the past few months, and discovered a pattern that determines the capability ceiling of Agents: beyond the model itself, beyond engineering and user experience polish, there's a critical piece of infrastructure — Agent Infra.
From memory, tool use, and task planning, to runtime or sandbox environments, to multi-Agent collaboration, even security and privacy mechanisms — every piece is indispensable. It's like when a new employee joins: you need to give them a computer, network cable, install Lark, DingTalk, before they can start working. Without proper infrastructure, both Agents and new hires are stuck. But challenges mean opportunities, which is why both startups and big tech have been doubling down on Agent Infra lately. Today we're delighted to welcome Alibaba Cloud VP and Wuying Business Unit President Xuqing, along with product manager Anchen, to chat with us about the past, present, and future of Agent Infra.
Alibaba Cloud's Wuying team launched AgentBay, a completely new experiment providing Agents with a complete runtime environment spanning cloud sandboxes to compute power and toolchains. We've prepared 20 questions, from technical principles to industry landscape, from entrepreneurial opportunities to career advice, hoping to help everyone build a clear mental framework for Agent Infra in this age of information overload.
Because we believe Agent Infra isn't just for technical leaders like Xuqing and Anchen — it will affect everyone's work and life like cloud computing did, and it will be an entrepreneurial opportunity for everyone. Let's start with a light question: what's the Agent product you use most in your daily life, and what do you use it for?
👦🏻 Anchen
The one I use most is called Onlook — it's an Agent tool that helps designers quickly design front-end applications. With Onlook, I can generate front-end interactive designs through natural language with one click. Another is an internally developed Agent from our group, built by the Taobao team, called Starflow. Its core capability is helping me quickly read foreign research papers.
👦🏻 Koji
What about you, Xuqing?
👦🏻 Xuqing
I use OneDay quite a bit — it's an Agent platform internally developed by our group. Externally, Cursor is also very good.
👦🏻 Koji
So Xuqing, you studied computer science, right? You wrote code yourself when you first entered the workforce?
👦🏻 Xuqing
Wrote code at Alibaba for several years.
👦🏻 Koji
What was the last year you wrote code?
👦🏻 Xuqing
Around 2017 or 2018.
👦🏻 Koji
Many technical leaders, like CTOs or the top technical executive, after moving into management focus more on strategy or architecture and stop writing code. But since Cursor appeared, many people have rediscovered the joy of coding.
👦🏻 Xuqing
Yes. Just a few days ago I visited Zhang Jianfeng, president of DAMO Academy, and he's also using Cursor to write code.
👦🏻 Koji
I'm curious — what is he writing with Cursor?
👦🏻 Xuqing
He demoed a small game for me, similar to Minesweeper in Windows.
The Agent Infra Landscape
👦🏻 Koji
Let's get into Agent Infra. First question: could you give everyone a primer on what Agent Infra is? How does it relate to traditional AI Infra? Where do they differ?
👦🏻 Anchen
Let me start with AI Infra, which people are more familiar with. Over the past few years when everyone was competing on large models, the focus was on token throughput, time-to-first-token latency, large-scale distributed training and inference efficiency, and cost. The problems behind these are what AI Infra addresses: how to achieve training, inference, and deployment through sufficiently good compute clusters, or platforms for publishing training jobs.
👦🏻 Xuqing
Like improving GPU efficiency.
👦🏻 Anchen
Right. The focus then was more on this part of Infra. It's not entirely IaaS — it could be PaaS too. Many enterprises use models for fine-tuning. There are many open-source models now. These are collectively called AI Infra. They ultimately focus on delivering Model Service.
But this year, people are paying more attention to building upper-layer applications. In the Agent Infra era, the previous Model Service becomes one component among many. To build an Agent application, beyond AI Infra, you need many other components — like Memory, task orchestration, tool use — and they are all important supporting pieces of Agent Infra.
👦🏻 Xuqing
I think Agent Infra has become an industry consensus because it shares certain conceptual foundations with traditional Infra. For example, compute: Agents need compute to execute code or browse web pages. Storage: like the long-term context memory often discussed in AI. Network: in Agent Infra, networks can be thought of as connecting to more capabilities. For instance, MCP lets Agents access more tools. Previously, large models might have focused just on reasoning, but now through these tools, they can actually take concrete actions.
👦🏻 Koji
Why has Agent Infra become a consensus this year?
👦🏻 Xuqing
At last year's World Artificial Intelligence Conference, people were still mostly discussing large models or chatbots. After this year's Spring Festival, Manus went viral, and the industry started discussing how to build universal Agents. At that time, Anchen, I, and some colleagues were invited to NVIDIA's GTC conference in the United States, where we saw more and more American companies working on Agents — almost every NVIDIA-invested company was talking about Agents. We felt we should empower companies like Manus by building a layer of Agent Infra, enabling them to build Agents more efficiently.
So we held an emergency overnight meeting to pivot our product direction. Drawing on Alibaba Cloud's capabilities in compute, storage, and networking infrastructure, we settled on AgentBay as our R&D focus. After about four or five months of team effort, we launched the product last week during the World Artificial Intelligence Conference, and the response from both internal and external stakeholders has been very positive.
👦🏻 Koji
We can dive deeper later into Wuying's transformation decision: because you're a very mature business unit with a core business that has strong growth, revenue, and profit, yet you were able to pivot overnight with an emergency meeting — that's a major strategic turn, and we'd like to discuss the management and strategic thinking behind it later.
But let's return to Agent Infra first. Could you two explain: what are the main components of Agent Infra generally? And what value does each provide for Agents?
👦🏻 安陈
When we discuss Agent design or deployment, there are some terms that sound familiar. Large models are like the CPU of an engine; long-term and short-term memory are like memory and storage, temporary cache and long-term storage. A hot concept this July was Context Engineering; there's also memory management, task orchestration, sandboxing, multi-Agent architecture, tool-use protocols, and so on. Previously no one explicitly defined what they looked like, but everyone was converging in this direction. Players gradually emerged in these segments, slowly forming industry-standard paradigms.
👦🏻 Koji
Let's inventory the key tracks in Agent Infra: first, memory; second, tool use; third, task planning; fourth, sandboxing; fifth, multi-Agent collaboration; sixth, security and privacy. We'll go through them one by one. Starting with memory. In the memory-related Agent Infra track, which companies and products are you two paying attention to?
👦🏻 旭卿
In memory, there's Memory0 and MemoryGPT — both are quite distinctive and go fairly deep. For tool use, there's E2B and BrowserBase, mainly solving Computer Use and Browser Use. Task planning depends more on the intelligence level of the model; everyone has some layout, so it's hard to say who's doing it best.
👦🏻 安陈
In my view, task planning is both orchestration and protocol. Most Agents today are multi-Agent architectures, so you need to define workflows for collaboration between multiple Agents. Common industry frameworks include LangGraph, and you can also use the OpenAI Agency SDK. We also have a purely self-developed Agent collaboration framework internally, though it hasn't been opened up yet. Another example is Google's A2A, which attempts to define a collaboration paradigm — it's more radical in implementation.
In security, the product we're building now is essentially security-related at its core. Code execution requires cloud-based isolated environments to ensure multi-tenant security isolation and zero local intrusion. Meanwhile, there are also companies in the industry focused on code security guardrails and sensitive data protection.
👦🏻 旭卿
Beyond security, identity management is also important. For example, during the Computer Use phase, when an Agent operates a computer and accesses different websites, it needs to automatically complete identity authentication. These kinds of problems are worth deep research, and we're also laying groundwork there.

👦🏻 Koji
Among these tracks, which one do you think is most important?
👦🏻 安陈
Actually, they're all important. For developers, every component is a hard requirement. Take sandboxing — users may not perceive it, but it's critical for secure code execution. Or memory: all Agents encounter context window limitations, so long-term memory is a universal need. While there's no clear priority, different components do have different sequences in terms of perception and implementation.
👦🏻 Koji
So right now, do Agent developers tend to build these components in-house or use third-party services?
👦🏻 旭卿
Before the Agent Infra concept emerged, almost every company wanted to vertically build everything in-house, but it's very difficult. It's like a new employee joining — if they have to handle their own desk, chair, computer, and network cables, efficiency will be very low. This is also a key constraint on Agent rapid development. If we can turn these components into systematic, easy-to-use services, Agent developers' lives become much easier.
👩🏻 Ronghui
It's like having to run and build the road at the same time — very hard. So when customers choose Agent services, what do they value most?
👦🏻 旭卿
Different companies focus on different things. For example, a major domestic model provider was also building its own Agent. Initially they built their own virtual code execution environment, Browser Use, and Mobile Use in-house. Later they discovered we had AgentBay and quickly switched over, because they needed 200,000 to 400,000 concurrent virtual machines — that's not a strength for a model company.
Another example is an automotive company. They're building an in-car intelligent agent that needs to remember the driver's long-term habits and questions, so they have very strong demands for long-term memory. We built a long-term memory system for them based on vector databases, ensuring users can consistently get a coherent interaction experience.
👩🏻 Ronghui
So your customization capabilities are quite strong?
👦🏻 旭卿
We provide general-purpose capabilities, but each company has different priorities when building Agents.
👦🏻 Koji
Among these modules, are there any that are very important in the long run but easily overlooked by developers today?
👦🏻 旭卿
Two: identity management and security. Many companies focus more on efficiency early on and neglect privacy and security. But in cloud computing, we've been trained from day one to consider data protection — it's in our DNA.
👦🏻 Koji
Got it — that's the muscle memory of running cloud services. Because once something goes wrong, the consequences are catastrophic. Let's talk about Alibaba Cloud's AgentBay then. In the year of Agents, Agent Infra has created massive startup opportunities. Why did you decide to build AgentBay?
👦🏻 旭卿
As large models and Agents evolved, we found they needed more "handy tools." Just as people use computers and browsers to improve efficiency, equipping intelligent agents with Computer Use, Mobile Use, and Browser Use makes them much more capable. We initially only planned to provide these foundational capabilities. But as we progressed, customer demands for long-term memory, identity management, and security became increasingly apparent, so we planned these capabilities too. This is the second phase of the product, more focused on core problems.
👦🏻 Koji
For example, companies like E2B, Browser Use, and LangGraph each focus only on their own vertical domain, while something like AWS's AgentCore and your AgentBay systematically cover the full chain. Being a "decathlete" where you have to do everything — do you feel your attention gets diluted? That you might not be able to compete with specialized startups in certain verticals? How do you think about this?
👦🏻 旭卿
We've always felt that toB business is a long-term track. A long-term track means relatively long-term investment in every domain. It's hard to say we can push every domain to the extreme in a short time. But we pursue maintaining above-average levels across the board, forming strong comprehensive capabilities overall. And we integrate the components systematically rather than in a fragmented way.
Wuying AgentBay
👦🏻 Koji
So if we're facing, say, our sandbox, right? Compared to E2B, what differences or advantages does our sandbox have?
👦🏻 安陈
To answer this question, let me first answer the previous one. Actually, the startups that are doing well in vertical domains today are more oriented toward top-tier developers. That is, these AI companies or agent companies already have the capability to stitch different components together and build complex application systems. But what cloud computing vendors want to do is essentially an inclusive business.
We want to turn these complex components into low-code platforms that even small and medium startups can easily get started with. So that's the positioning difference between us and them. Returning to your question, I think our advantages are roughly threefold:
First, ease of use. We designed from the start for small and medium developers, so the product supported MCP out of the gate, right? Making it easier for developers to get started.
Second, completeness. Our sandbox isn't just a code environment — it includes mobile environments, browser environments. We've also connected the unified persistence system underlying the agent service, and will add memory context and so on in the future. So for customers, this isn't just a vertical product, but more like a complete agent-building platform.
Third, forward-looking. We pay more attention to the long-term development direction of agents. For example, some top-tier developers may only want you to give them an "atomic" sandbox component for them to perceive and control themselves. But SMB developers don't want to manage various agents inside the sandbox themselves — they want the platform itself to provide stronger perception and control capabilities.
For example: if I just give you a browser container and have you schedule it yourself, the simplest way is to install an open-source provider and do visual manipulation through the DOM tree. But in a browser environment, relying solely on the DOM tree isn't enough — for iframes or video content, the model can't perceive anything through text input alone. At this point, you need to combine multimodal models to truly understand what's happening on screen.
So as a provider of browser user services, we must research and assemble multimodal models that can both perceive the browser screen and output multimodally. Only then can we truly help customers complete the environment's driving and perception capabilities.
👩🏻 Ronghui
Right, so strategically you emphasize inclusiveness more, valuing SMB customers more.
👦🏻 旭卿
Yes. For example, all our capabilities are exposed through APIs — developers only need simple API calls to obtain corresponding functionality.
👩🏻 Ronghui
That's very much in Alibaba's style — making it easy to build agents for everyone under the sun.
👦🏻 Koji
We know AWS also launched AgentCall. How do you see the differences in positioning or priority between AgentCall and AgentBay? Could you walk us through the similarities and differences?
👦🏻 Xuqing
We launched AgentBay around April 7th at our summit. At that time it was still more of a concept and product design prototype, and we opened up some simple capabilities for external customers to try. By the World AI Conference, we officially released the commercial version, allowing small and medium-sized developers to use the full product directly on our website.
Around the same time, we saw AWS introduce AgentCall. After their launch, we did a comparative analysis and found that when looking at Agent Infra, everyone's thinking is actually quite aligned. Whether it's product capabilities, API definitions, or overall layout, the similarity is quite high.
👦🏻 Koji
Do you get the sense, though, that among AWS, Volcano Engine, or other cloud providers, someone is investing the most heavily in Agent Infra? Like in terms of funding, team size, even the commitment of the top executive — are there clear differences?
👦🏻 Xuqing
This domain represents a critical application direction for AI across the entire industry. So I believe every company takes it very seriously. Taking Alibaba Cloud as an example, our CEO Eddie Wu is actually very focused on AgentBay and Agent Infra. At this year's all-hands kickoff meeting, he specifically highlighted Agent Infra. Before our product launch, we also gave him a fresh round of updates, and he was very satisfied with our progress.
In this space, leadership is also quite willing to invest — they recently added headcount, budget, and resources for us. You could say that from the CEO down through every business unit to the frontline teams, everyone values this direction and recognizes it as the right strategy.
👦🏻 Koji
Why such heavy emphasis? Is this a defensive play or an offensive one?
👦🏻 Xuqing
I believe from the CEO's perspective, this isn't viewed purely as a business yet, but rather as infrastructure for AI applications. If a cloud provider has the resources and capital but doesn't build these capabilities, it would actually slow down AI progress.
Also, while leadership has added people and resources for us, they've never required us to hit specific revenue targets. Instead, they care more about how many customers are actually using the product.
👦🏻 Koji
That's interesting, because our next question was going to be: with all this added money, headcount, and resources, did OKRs go up? Did KPIs go up?
👦🏻 Xuqing
That definitely didn't happen, because at this stage our assessment of AI is still relatively early-phase. We still need more companies like Alibaba Cloud to participate and build out the entire ecosystem together.
👦🏻 Koji
So the main metric you're looking at is whether more people are using your service, not how much they're paying for it at this point?
👦🏻 Xuqing
Revenue is not the priority right now, definitely not.
👦🏻 Koji
So I imagine you probably don't need to help Alibaba Cloud sell cloud services either, right? That's not a metric?
👦🏻 Xuqing
There's no explicit target there. We prioritize market share.
👩🏻 Ronghui
So what position does AgentBay hold within Alibaba Cloud's AI strategy?
👦🏻 Xuqing
AgentBay — over the past two years we've seen everyone competing on foundation models, but how to actually make good use of AI? Agent is a critical direction. So within the broader AI picture, we play a connective role. On one hand, we expose model capabilities outward through Agent Infra. On the other hand, we'll likely build some Agent development frameworks to make it simpler for more users to develop Agents.

👦🏻 Koji
Last week we attended the livestream for Zhou Hongyi's new product launch: the Nami Intelligent Agent Swarm. He talked about a choice they made when building Agents — they didn't go with cloud sandbox, but chose local instead. He made an important point: he feels cloud sandboxes aren't secure. Users putting their passwords, personal authentication info in the cloud could lead to serious problems, so they chose local. This is a very firm strategic choice from an old hand in security.
So I wanted to ask you two, because AgentBay's sandbox is cloud-based, and you've also mentioned that whether at Alibaba Cloud or in previous roles, you've always prioritized security highest. Could you explain how you convince developers that AgentBay's cloud sandbox is secure?
👦🏻 Xuqing
The essence is whether the cloud can provide a more secure environment. I started doing cloud computing research around 2005 and joined Alibaba Cloud in 2014. My earliest understanding of cloud was that maybe putting data in self-built data centers would be more secure. But in my second year at Alibaba Cloud, my perspective changed. At that time, a game company with its own data center launched a new game and got taken down by a competitor who bought 5GB of traffic with Bitcoin. But on Alibaba Cloud, DDoS resistance is extremely strong, so many game companies migrated to the cloud. The biggest difference with public cloud is continuous iteration.
We know any software and system has vulnerabilities, but cloud providers — we have a security team of two to three thousand people guarding data security in real-time. A company building its own data center, first, can't hire security engineers this good, and second, the cost is extremely high. My decade-plus of cloud computing experience has deeply convinced me that public cloud provides far stronger security capabilities than enterprise self-build.
👦🏻 Koji
OK, I understand the difference here is that using cloud sandbox versus local sandbox — local sandbox isn't local cloud, it's literally the user's own computer.
👦🏻 Xuqing
Right. First, from a capability perspective, if you use local sandbox, like Manus founder Red Xiao talked about a few days ago, being able to spin up 100 sub-agents simultaneously to do Wide Research. This kind of workload is very difficult to complete on a local computer — it can only be done in the cloud. From a capability standpoint, local execution is increasingly unrealistic.
Additionally, in the cloud, from the overall sandbox design to the security architecture at the Agent Infra layer, we provide end-to-end data protection. It's like keeping money at home versus putting it in a bank — the security is equivalent. Data security is the lifeline for any cloud provider; everyone invests massive human and material resources to ensure data isn't leaked or stolen.
👦🏻 Koji
But this is also a strategic choice. Some people feel local has another advantage: you don't need to provide new usernames and passwords when logging in, because the browser already stores cookies. How do you convince — and I believe there are agent developers who will have preferences here — how do you convince them?
👦🏻 Xuqing
Like you just mentioned, browsers have many cookies storing lots of usernames and passwords. In that case, what you're choosing to trust is the OS vendor and the browser vendor. Because the OS vendor and browser vendor can definitely access that data.
But why do we trust today that Microsoft or Google won't touch your stuff, and why can't we choose to trust a cloud provider? It's just that you feel data stored locally is more secure, but in reality local computers may have trojans and viruses that more easily cause leaks. When you entrust this data to a third-party secure vault, that environment is relatively more singular and pure.
👦🏻 Koji
Another argument is that whatever code runs in a cloud sandbox, if it crashes it's the cloud's problem; local is safe.
👦🏻 Xuqing
There was a story I thought was a joke, but learned a few days ago it's real: a developer had a model actively delete a directory in a local sandbox, and it actually deleted that directory on the computer. But this isn't a problem in a cloud sandbox.
👦🏻 Koji
Because in a cloud sandbox, deletion can be recovered — it's virtual, it won't harm your local personal files or private information. Then for example, we mentioned Manus's Wide Research release last week, and we're seeing a trend: last year people were wondering if token consumption was about to peak, if NVIDIA's cards would stop selling, and we did see prices dropping. But this year with agents, it's not the case — as capabilities improve, token consumption is increasing tenfold or even hundredfold. Have you seen customers with high concurrency demands? What's the highest level customers have required from you?
👦🏻 Xuqing
A domestic agent company, during normal daytime hours they need to handle roughly 200,000 concurrent sessions. That number surprised us quite a bit. But analyzing it, it makes sense — because if it's an app or application, the user count determines the concurrency ceiling.
👦🏻 Koji
Is this 200,000 meaning 200,000 users simultaneously issuing agent tasks, or is it say 10,000 agent tasks behind the scenes, each spawning 20 sub-tasks?
👦🏻 Xuqing
These are concurrent sub-tasks.
👦🏻 Koji
Then it really does multiply rapidly. Like Manus's Wide Research case is one main task immediately spinning up 100, 200 sub-tasks.
👦🏻 Xuqing
We believe this is an important direction for future general-purpose scenario development.
👦🏻 Koji
That sounds quite profitable.
👦🏻 Xuqing
We haven't been thinking about profitability, but rather how to better meet the needs of companies like Manus with our product capabilities.
👦🏻 Koji
So there's no very clear pricing model conclusion at this point?
👦🏻 Xuqing
We've looked at this space, and when building the product, at this stage it's still about delivering the capabilities and meeting all these agent development needs. But long-term, it's definitely still a commercial activity.
👩🏻 Ronghui
What new business models might emerge in this space in the future? How might pricing methods differ from before?
👦🏻 Anchen
Currently as a cloud computing company, we're still primarily in a compute-power sales model. For example, if you're a manager and need to spin up 100 concurrent instances, I charge you for 100 sandboxes. But in the future, it may shift more toward application-layer pricing.
For example, when a customer has a need, what we provide might not just be the compute environment, but also bundled multimodal model inference costs, state consistency management costs, and so on. In this scenario, compute power might represent only a small fraction of the cost, and the pricing model becomes much more service-oriented.
👩🏻 Ronghui
What does your customer demand curve look like right now?
👦🏻 Anchen
It's a very exponential curve.
👦🏻 Xuqing
You can see user numbers multiplying, with tons of new users coming in every day, and very high concurrency. So this morning before coming here, we even had a meeting to discuss how to achieve global resource scheduling.
Opportunities and Choices
👦🏻 Koji
Since many listeners of the Crossing podcast are entrepreneurs or investors. I think hearing you talk about exponential growth up to this point, everyone probably wants to ask: beyond general-purpose agents, what vertical domains are growing very fast, whether B2B or B2C? What are your observations on these areas? Because for everyone here, this means investment opportunities or entrepreneurial opportunities.
👩🏻 Ronghui
Right, the "pick-and-shovel" companies definitely feel this most acutely.
👦🏻 Xuqing
We have a saying from our cloud computing days: you can very clearly feel the changes across every industry in the entire sector.
👦🏻 Koji
Yes.
👦🏻 Anchen
What we're mainly serving now are general-purpose agents and coding agents, which everyone as consumers can also feel. The commercialization model for coding agents is relatively mature, with the broadest scope. The second category is general-purpose agents—since Manus's release, all manufacturers have been making transformations, which everyone can feel.
Beyond that, based on the general-purpose agent framework, as infrastructure gradually matures, we're seeing many traditional applications also undergoing agent-based transformation, mostly inside enterprises, which consumers may not perceive.
For example, some e-commerce companies have large amounts of automated work in their traditional operations, like IPA tasks. Operations staff need to do cross-platform price comparisons, product listings, main image design, etc.—previously large amounts of repetitive workflows, now all being transformed with agents.
Including many major clients we're now working with—not integrating agents to provide external services, but doing internal OA automation, or automation of repetitive work. This manifests as automated operations, repetitive finance work, internal customer service, and so on. This is subtle, mainly serving to improve efficiency within large enterprises, rather than being directly customer-facing.
👦🏻 Koji
Got it, what about customer-facing ones?
👦🏻 Anchen
Customer-facing, for example HR agents (human resources-related), and in finance there are some doing investment recommendations, many people are working on this.
👦🏻 Koji
Right, tons of people are doing this. Because it's close to the money, investment is easier to close the loop—finance is all numbers.
👦🏻 Anchen
Right, I also made one myself not long ago.
👦🏻 Koji
Hahaha, helping you trade stocks?
👦🏻 Anchen
Yes, doing investments. Everyone knows news and information are crucial, so you need to broadly acquire data and have AI help you analyze it—many people are doing this. Quant is the place closest to money, so there are many agents related to investment recommendations.
Additionally, there's healthcare and medical-related. The US medical technology index is doing very well right now, largely because large models can create more value for them.
👦🏻 Koji
Our podcast has talked to many entrepreneurs before, but relatively few senior executives from major tech companies. I myself have many questions about doing business and management at a big company, and hope to ask Xuqing about them today.
First, the Wuying division makes cloud PCs, and from previous conversations with you two I also understand it's growing very fast and has decent profitability. In this situation, I understand that making a transformation requires conviction, and also has significant opportunity cost. Because the new direction may need a long time and much cultivation before seeing returns, unlike the original business where slight expansion or globalization can quickly show growth, making it easier to deliver results in year-end reports.
So I want to ask, how was such a major decision made?
👦🏻 Xuqing
My view on B2B business is that any B2B business is long-term. Unlike selling phones, where good design and popularity can lead to immediate blockbuster sales. B2B business is always long-cycle investment with slower results.
I remember very clearly, before the 2016 Spring Festival, Teacher Ma came to Alibaba Cloud and said: investing in any B2B business at Alibaba has a 10-year investment cycle. For example, Taobao started in 2003-2004 and didn't go public until 2014; Alipay started in 2006 and only became relatively mature by 2016; Alibaba Cloud started in 2009 and only grew very rapidly around 2019.
From this you can see that in Alibaba's B2B business investments, within ten years there's basically no requirement to achieve a certain scale or make a certain amount of money. Instead, it's about whether you can deeply cultivate the赛道 and create value, especially long-term value. So in everything I do, I think in ten-year units. After Eddie Wu came last year, he upgraded Wuying to a first-level division, treating it as a very important赛道. Last year we made Wuying and large models the two important strategic new products of Alibaba Cloud Intelligence Group.
👦🏻 Koji
But this year you adjusted the cloud PC赛道 that was considered very important last year?
👦🏻 Xuqing
No, this may be a misunderstanding. Wuying's growth rate in recent years is still very high, consistently maintaining triple-digit growth. From a business perspective, last year we promoted it as the most important strategic product in terminal cloud computing. In this process, large models and intelligent agents were also happening, and we were looking at how Wuying could integrate with agents and AI to provide energy加持 for the long-term future. How to make cloud PCs become good computers, phones, or code execution environments for AI, agents, and models in the large model era? We started laying out based on this thinking.
The real determination to do it came after ChatGPT. After thinking it through clearly, we positioned AgentBay as a sibling relationship to Wuying, or an extension in product and technical capabilities. Of course, it's simple to say but hard to do. You need resources, people, more manpower invested in this. The past six months I've been quite conflicted—on one hand ensuring this side's business maintains high growth, on the other hand ensuring we don't miss the new赛道 and can execute efficiently and effectively.
👦🏻 Koji
Right, this is quite difficult.
👦🏻 Xuqing
It really is quite difficult. In internal coordination, some colleagues really want to do AI and volunteer to join; others feel they're doing very well now with important work, but get forcibly transferred over. There's a lot of trade-offs and communication in this.
Management and Career
👦🏻 Koji
Because I understand there's actually a very difficult point here. AgentBay is a very bright opportunity for the future—could there be colleagues still working on Wuying who feel somewhat disappointed, not having been transferred to AgentBay?
👦🏻 Xuqing
I believe there definitely are. But overall, both sides are strategic businesses. It's just that AgentBay may be more long-term, while Wuying has already landed and maintains very good growth momentum. I believe there's a sense of achievement in working on both sides.
👦🏻 Anchen
Let me make a provocative statement. Hahaha, the boss may find this inconvenient to say, so I'll be slightly more radical. I actually think that what Wuying has been doing with cloud PCs, our past concept was called "endpoint compute power to the cloud." Because we believe individual users' compute requirements can't be met by traditional PCs. So for a long time we served specialized industries.
For example, in recent years many individual designers use SD, use Flux, and some doing simulation need highly elastic, high-GPU-resource compute, so they go to the cloud to use cloud PCs. There are also some security clients. But the general public hasn't really perceived this.
This year is different. This year the concept of computers is being reconstructed. Many people say this is the year of the "super individual," the year of "personal cloud computing." Individuals' compute requirements can't be fulfilled by simply getting a computer. For example, if I need to run 100 environments now, it's completely impossible locally. So in this era, cloud PCs will truly realize endpoint commercial scenarios. We believe the general public will all need cloud PCs—this is what we've been doing for years.
AgentBay, while serving some B2B clients, also serves cloud PCs. We're now building a logical cloud PC:
First, it can be persistent, with state and data residing on it;
Second, it's highly concurrent and highly elastic, able to mirror out 100 cloud instances running concurrently during large-scale task scheduling;
Third, it can roam anywhere, across multiple devices. Use it on phone today, Pad tomorrow, computer the day after—of course we also have our own hardware;
Finally, it can be driven by natural language. This is something many traditional computers can't do. But AgentBay is fundamentally about letting AI use a computer.
So we believe future computers will first be cloud-based, second be highly concurrent, third be accessible to everyone—from 80-year-old elders to 7-year-old children, all can drive computers through natural language. At this point, cloud PCs have arrived at their historical mission. So many of our colleagues in the traditional business are still fighting to upgrade and transform the business.
👦🏻 Xuqing
Right, essentially it's still agent-driven computers. Whether we're sleeping or chatting, that agent-driven cloud PC is actually doing things for you. We've also mentioned the concept of "digital employees," hoping that every employee has a digital twin in the cloud, using intelligence enhancement and cloud PC capabilities to parallel-complete many complex tasks. A few days ago we also mentioned "silicon-carbon symbiosis." Meaning between digital people in the cloud (silicon-based) and real people, there can be much coordination and interaction.
👦🏻 Koji
I understand you two have attended many exhibitions and tech summits in your careers. But it sounds like this NVIDIA GTC brought you very strong shock, so much so that you wanted to set strategic plans for the next few years that very night. Have you had this feeling before in your previous career?
👦🏻 Xuqing
In 2016 or 2017, I went to the US for an exhibition, the specific name I forget. That year containers were extremely hot, every company was talking about containers, and there were many container-related startups.
But containers actually running on physical machines or virtual machines would suffer performance losses. Because at the filesystem level it used OverlayFS—convenient but performance degraded; networking was also the most primitive virtual network, using software virtualization with performance limitations. These would seriously affect large-scale container adoption.
After returning we laid out a new product, which later became called "X-Dragon" servers, officially released in 2017, and later the entire industry followed, using our standards to make DPUs and bare-metal virtualization.
👦🏻 Koji
We previously told friends that Xuqing was coming on Crossing, and a friend said this is the "god-level boss of elastic computing." At the time I thought of another angle: does a god-level boss need to transform?
Because we recently published an episode called "Programmers at the Crossroads in the AI Era." There was fresh data from the United States: this year, computer science graduates had a 6.2% unemployment rate, while arts students were at 3% — computer science unemployment was double that of arts.
So when I saw "the god-level boss of elastic computing," on one hand I was thrilled to welcome such a heavyweight guest. But on the other hand, I also wondered: today it's not just fresh graduates — many people at mid-career are suddenly facing such massive AI disruption. Do they panic? Do you ever feel like, you finally became this god-level boss, and now here comes another giant technology wave that will change everything? How do you understand and face that?
👦🏻 Xuqing
Looking back at my ten-plus, twenty-year career, I think I've been someone who likes to think, who thinks diligently.
Before coming to Alibaba, I worked at Intel for nine years (three as an intern, six full-time). From 2008 to 2014. I was probably one of the few students at Intel who interned for three years, because I was still doing my PhD then. At that time, the whole world was discussing what cloud computing was and where it should go.
I was fortunate to join Intel's system virtualization team, mainly working on open-source technologies like KVM and Kubernetes, which were later used in cloud computing. Another classmate who joined with me felt this direction was too niche and might not lead to good job prospects. At the time, people researching virtualization globally were basically just at VMware, Microsoft, Cambridge, and Stanford — maybe fewer than 100 technical people total. Many thought this was a niche track.
But after I joined, I became somewhat obsessed. I already liked CPU, operating systems, these low-level technologies, and I discovered virtualization let you go even deeper across more layers — hugely attractive. Although back then no one had a concept of what technologies cloud computing would need, I felt such profound technology must have important applications.
Before I even graduated, AWS was showing its potential in 2006 and 2007. Then the whole world came to our (Intel) team to poach people. Our team had about twenty-seven or twenty-eight people, and more than ten went to American companies, all getting US offers and relocating directly. Because at that time, people familiar with core cloud computing technology globally probably numbered fewer than 50, and in retrospect maybe not more than 100.
👦🏻 Koji
Sounds kind of like the people doing large model research today.
👦🏻 旭卿
Right, at that point in time I felt quite lucky to participate in the cloud computing wave. Around 2014 I went to Alibaba. I told my leader then, how can we build a computing platform serving tens of millions of customers? Everyone thought it was a pipe dream, but we were very determined — if certain technical architectures weren't conducive to scaling to tens of millions of customers, we had to scrap them and start over. In 2014 we chose correctly, because in 2015 mobile internet exploded, and after all kinds of apps took off, when those customers came onboard our product was already ready.
By 2016, we optimized virtual machine performance to within three to five percentage points of physical machines, and everyone thought that was the industry limit, already a benchmark. But I was thinking then, sacrificing some CPU resources to improve performance — that's not the right direction. In 2016 I thought about these questions every day. Later at a trade show, I saw container and other new technologies that seemed performance-lossless, but still had lots of underlying overhead. So how to improve computing performance through software coordination, chip design, and these areas? So in 2016 we laid out X-Dragon, which became an industry benchmark product in one move.
Almost all cloud computing companies now use underlying technical architectures not much different from X-Dragon, because after we released it the whole world followed. Speaking of AI, I started laying out what you'd call AI infrastructure in 2015. The reason was that Alibaba was using NVIDIA K2 graphics cards internally — pretty poor computing power by today's standards — but we were using them for machine learning, doing a project called "Pailitao," where Taobao would identify photos and recommend similar products.
Seeing this direction, I felt Alibaba would have this demand in the future, and external companies would too, so I started laying out AI infrastructure in 2015. By end of 2016 the product was already done, just in time for the AI wave explosion led by deep learning in computer vision, speech, and so on. In this wave we served over 80% of Chinese tech companies' AI computing demand — this was because we laid out the product in 2015. In 2017 we made X-Dragon, while also laying out AI infrastructure for future large-scale parameters, which later became GPU supercomputer clusters.
Over these years I feel overall you still need to think for the future, especially after moving into tech management — you find you need to make some predictions and preparations for what might happen in two or three years, rather than waiting for the wave to hit and then defending, which would be too late.
👦🏻 Koji
So now looking at agents, among the things you're doing for the future, is there anything you worry you might bet wrong on?
👦🏻 旭卿
I don't worry about any particular thing being wrong, but rather whether there's some piece in the overall layout that I didn't think of — this is what I think about every day.
👦🏻 Koji
You don't know what you don't know.
👦🏻 旭卿
Right, that's the scariest.
👩🏻 Ronghui
I see. Now it feels like a major trend is "you need to know a bit of everything." Especially one-person companies, where one person can do many things. Would you advise a junior engineer, or someone with a few years of experience, to strengthen capabilities in a specific domain, or seeing current trends, to know a bit of everything?
👦🏻 旭卿
I think it's different at different stages. If you're an engineer just entering the workforce, you still need to focus on a specific domain and go deep. Breadth of knowledge is of course important, but without deep accumulation in some domain, future development will be limited.
It's not about how much you know, because now large models know everything, they can answer any question. But to do very specialized technology requiring deep research, large models aren't yet good at it, especially when combined with human thinking. Many things aren't solvable just by large models having broad knowledge.
For engineers, my suggestion is like what I did back then — research virtualization technology deeply, then combine it with the business you're working on to bring out its value. This way, starting from one technology, you drive understanding of surrounding technologies, the product becomes competitive, and the business grows better.
👩🏻 Ronghui
The bar for people keeps getting higher.

👦🏻 Koji
I'm quite curious — listening to Xuqing, whether during the internship or in 2015, 2016, first seeing cloud computing, then elastic computing, and now the growth of Agent Infra, many of your predictions have been validated as correct, even stepping on the biggest waves in the tech world. Have you summarized why you could make these successful predictions? What qualities, habits, or ways of thinking does this require?
👦🏻 旭卿
It's really about finding future clues in what's happening today. For example, at the 2016 Double Eleven summary meeting, our CTO at the time asked whether virtualization technology with only 3% to 5% performance loss could achieve zero loss. If I hadn't attended that meeting, or had ignored that statement thinking it was impossible, already极致 — then there might have been no progress. But at the time I felt I was the number one technical person, and the boss raised a seemingly impossible requirement — could I think differently, would this happen in the future?
Later at a trade show, I found related technologies were developing rapidly. Putting these clues together, I proposed a likely future direction: hardware-software co-design, deep collaborative optimization. In our domain, could hardware-software collaboration bring revolutionary change? So I thought every day: to achieve this goal, which technologies aren't sufficient today? What can't be done with existing chips and system software? Then do we need to lay out a chip? Lay out certain system technology R&D?
When these became clear, I proposed making the X-Dragon chip and convinced management to invest. At that time no one at Alibaba had made chips before, but I felt I had to convince them, because if the opportunity appeared and we didn't invest, and it was validated three to five years later, the company might miss a strategic transformation opportunity. So basically I always look for clues from today's problems and difficulties, seeing if they might be solvable in the future.
👦🏻 Koji
I'm reminded of an article we both like, Paul Graham's How to Do Great Work. It says you need to find that thing you do effortlessly that others find hard. That thing is likely your starting point for doing what others can't. Do you have this feeling? Along the way, outsiders see many things you've done as very difficult, huge challenges. But doing them yourself, you feel quite happy, quite at ease — like you've found that thing you do effortlessly that others find hard.
👦🏻 旭卿
Quite the opposite. After joining Intel, in that team, I watched the technical experts around me. In meetings they all spoke Chinese, every sentence I understood the literal meaning, but I didn't know what they were discussing. At the time I felt I needed half a year to understand what they were doing, what they were discussing. If I didn't understand and gave up, that wouldn't be my personality.
👦🏻 Koji
How do you think this personality of yours, or this desire and habit of learning, was formed?
👩🏻 Ronghui
Been a good student since childhood.
👦🏻 旭卿
I think this might be related to personality, and personality might be innate. Call it being competitive or whatever, but overall I like seeing difficult things, not easy things. I prefer things with more challenge.
👦🏻 Koji
Knowing there are tigers on the mountain, still heading toward the tiger mountain. So do you feel the challenges of doing Agent Infra today compared to the past — what level of difficulty is this?
👦🏻 旭卿
Laying out Agent Infra actually requires me to learn many new things. I often need to find product managers, find tech R&D leaders to discuss various things, because it's somewhat different from what I learned and was familiar with before. For me personally, this is quite challenging. I need to understand it to communicate with them as equals, otherwise I wouldn't even know if I'm being fooled.
👩🏻 Ronghui
When you see a future direction and want the company to invest in developing it, but it might be unknown, how do you convince the boss?
👦🏻 旭卿
This is an interesting story. For example, when I wanted to build X-Dragon — internally we called it the MOC card (the industry now calls it DPU). I had only been at Alibaba for two years. But many of my previous decisions had been proven right, so people trusted my judgment. Then in 2016, I suddenly told my boss I wanted to build chips. That's a challenge for any boss. I would wait for my leader to arrive at the office almost every morning and explain how important this was. I did this for about half a month before I finally convinced him. It was later proven correct too. When we launched it at the 2017 Apsara Conference, cloud computing companies worldwide were surprised — they didn't realize you could do it this way. Previously, everyone thought the ceiling of a virtual machine was the physical machine. Now, the ceiling of a virtual machine is unlimited. When this concept gained recognition, the team felt a real sense of accomplishment.
👩🏻 Ronghui
In the workplace, people on the front lines may not be blind to the future — the challenge is how to convince the boss to invest resources.
👦🏻 Xuqing
So I hope that when colleagues on the team face difficulties, they can rise to the challenge and persuade me to make these investments.
👦🏻 Koji
Just now Xuqing mentioned that you've gotten more headcount and resources this year. Want to do some recruiting here?
👦🏻 Xuqing
Of course. We really hope more people who believe in this direction will join us — whether fresh graduates or senior developers with experience in AI — to build the future Agent Infra platform together, and make it more efficient for more enterprises to develop their own agents.
👦🏻 Koji
Alright, thank you both for today. Thanks for your time.
🚥
