"The 100 Million Token Club" Is Packed — AI Is Running Out of Fuel | A Conversation with Wenyuan Yu: Technical Lead of Alibaba Cloud Bailian
Agent Ignites the Compute Explosion
Agent Ignites a Compute Explosion

👦🏻 Podcast Interview: Koji
🥷 Editing: Crossing
🧑🎨 Layout: NCon

🚥 If you've been using Claude Code, OpenClaw, or various Agents lately, you've probably noticed something: models are getting better, SOTA results are exciting, but Tokens are a wake-up call — expensive, and never enough.
Someone even created a group for power users with a ridiculous threshold: burn through 100 million Tokens a day to join the "100M TOKEN Club." What's more ridiculous is that this bar is getting "not high enough" — because more and more people are pushing AI from chat tools into real productivity workflows.

For this episode of Crossing, our guest is Wenyuan Yu, technical lead of Alibaba Cloud's Bailian. Yu occupies a rare vantage point: how compute demand is skyrocketing, which scenarios are devouring Tokens, how the cloud paradigm is being rewritten by Agent, and the real engineering challenges behind "no amount of GPUs is enough."
We also discuss why exploding Token consumption isn't a temporary bubble but a phase-defining signal; which enterprises should actually build their own infra; and why AI coding's rise makes vibe coding more dangerous, not less.
If you care about how AI enters production, stabilizes, scales, and where the next wave of opportunity will emerge, this episode is worth your time.


Listen on WeChat:
Listen on Xiaoyuzhou:

🎬 Video podcast now live on Koji's WeChat Channels, Xiaohongshu, Bilibili, YouTube, and other platforms
Given the full interview's length, here's the table of contents:
🟢 Rapid Fire
Age, alma mater, MBTI and zodiac, one-sentence intro to Bailian, career history.
🟢 Token Nuke: The Compute Explosion Ignited by Agent
Claude Code and OpenClaw sweep the globe — behind this isn't just a tool going viral, but a fundamental shift in how compute gets consumed.
🟢 Every GPU Second Must Work
There's a CEO with "the most aggressive compute investment" — and it's still not enough. This hunger for compute has never existed in cloud computing history.
🟢 Build In-House or Cloud? I'll Make a "Hot Take"
Cost control, data security, flexibility — these three reasons enterprises cite for building their own GPU clusters are, Yu argues, precisely why they should use MaaS instead.
🟢 Unpopular Opinion: Don't Let AI Write Too Much Code for You
His advice to CS students is "use AI less" — isn't that contradictory?
🟢 Most Counterintuitive Prediction: OS Developers Get Replaced by AI First
Everyone thinks frontend engineers are most at risk — Yu says the opposite.
🟢 Compute Is Like Oil: But Today's Key Isn't Just Compute
China's cloud reference architecture has historically come from the United States — but this time, even the U.S. doesn't have answers yet.
🟢 Future Infrastructure Stack: Water, Electricity, Coal — and Models
AI will become a commodity like water, electricity, and coal — but Yu says: yes, yet it won't become that kind of "plug in for 220 volts" standardized infrastructure.

Rapid Fire
👦🏻 Koji
Let's start with rapid fire. Wenyuan, your age?
👨🏻💻 Wenyuan Yu
👦🏻 Koji
Where did you study?
👨🏻💻 Wenyuan Yu
Undergrad at Peking University, PhD at University of Edinburgh.
👦🏻 Koji
MBTI and zodiac?
👨🏻💻 Wenyuan Yu
INTP, Pisces.
👦🏻 Koji
What were you doing before Bailian?
👨🏻💻 Wenyuan Yu
In 2018 I founded a small company that got acquired by Alibaba. At the time I was working on graph computing. Later I did systems research at DAMO Academy and Tongyi Lab.
Token Nuke: The Compute Explosion Ignited by Agent
👦🏻 Koji
Over the past two months, Claude Code and OpenClaw have swept the globe. What impact has this had on your work?
👨🏻💻 Wenyuan Yu
The most direct effect is rapid Token growth, basically doubling month over month. And these are all high-quality, SOTA coding models. People no longer treat AI as a chatbot or casual conversation tool — they're integrating it into productivity scenarios.
This is an absolutely massive Token burn, and it's only just beginning. We believe this growth rate will quickly reach staggering levels.
👦🏻 Koji
Only just beginning?
👨🏻💻 Wenyuan Yu
Really just the beginning.
👦🏻 Koji
On one hand there's Claude Code, on the other OpenClaw, creating nuclear-level Token consumption explosions, with global compute shortages. Are there other factors?
👨🏻💻 Wenyuan Yu
Today, AI is fundamentally changing how people consume compute. I can't predict what the next explosive scenario will be in the near term, but I'm absolutely certain that in three to five years, much of today's labor-intensive work will be done by AI.
Second is cloud computing itself. What do data centers and scheduling systems look like today? How do people consume compute, storage, networking? In three to five years, it'll look completely different.
👦🏻 Koji
So you believe cloud computing's paradigm will transform dramatically?
👨🏻💻 Wenyuan Yu
Dramatically.
👦🏻 Koji
Will cloud providers get reshuffled?
👨🏻💻 Wenyuan Yu
We're already seeing signs of reshuffling.
The definition of "what makes a good cloud" is changing today. Why are new clouds emerging vigorously? What will the future cloud look like? We're at a crossroads of development.
👦🏻 Koji
For a while, the consensus was that global cloud provider格局 was settled — just those few giants in China and the U.S. But with Neocloud's emergence, do you think this is a flash in the pan, with the market eventually captured by cloud giants? Or do these new companies have a shot at the top tier?
👨🏻💻 Wenyuan Yu
That's hard for me to judge. But every cloud provider is also transforming, revolutionizing themselves.

Alibaba Cloud is China's largest cloud provider, and we're also thinking: will our future cloud users no longer be humans, but Agents? What compute, storage, and networking do Agents need from the cloud? What databases and compute power do they require? How do we meet their needs? Every provider is preparing for this transformation.
Every GPU Second Must Work
👦🏻 Koji
What are your top priorities at Alibaba Cloud?
👨🏻💻 Wenyuan Yu
Stability is definitely our top concern. Security is also extremely important, but what we see changing daily is: user volume, Token volume, and massive compute demand growth.
Our Qwen 3.5 model launched on New Year's Eve, and within just over two weeks, its peak (requests per minute) reached heights no text model in history had ever seen.
We have the most aggressively compute-invested CEO, and he still feels it's not enough. Because we have massive model R&D and customer service demands, our Token growth pressure remains enormous.
👦🏻 Koji
When did this explosive growth start?
👨🏻💻 Wenyuan Yu
My sense is it never stopped from day one of Bailian's launch.
👦🏻 Koji
So you don't feel OpenClaw or Claude Code accelerated this growth?
👨🏻💻 Wenyuan Yu
In Agent scenarios, they absolutely accelerated growth. But historically, we've been through many such explosions — like with certain video generation models, once they cross the threshold from demo to market-viable, they rapidly bring wave after wave of growth.
👦🏻 Koji
In the midst of such dramatic change, is it hard to pursue stability?
👨🏻💻 Wenyuan Yu
Extremely hard. And on top of seeking stability, we also want to make sure we're fully utilizing our compute.
Beyond security and stability, there's another major constraint we face with such rapid growth — GPU supply is extremely limited. Algorithm teams are all competing for them, saying they need GPUs for training, wanting to deploy stronger models for certain scenarios, demanding better service quality. There's enormous systems and engineering challenges in all of this.
We have a very important mission: not a single second of any GPU should sit idle. We want every GPU to operate at maximum capacity. We want to put all of Alibaba Cloud's GPU compute resources to work — whether that's 1,000 cards, 10,000 cards, or 1 million cards — so that every user can enjoy extreme elasticity and stability, feeling like they have access to China's largest compute cluster behind them, callable through a single API.
👦🏻 Koji
There's been a WeChat group recently called the "100 Million Token Club," where you need to consume 100 million tokens daily to join. You mentioned earlier that we shouldn't just focus on quantity but also quality. Could you elaborate?
👨🏻💻 Wenyuan Yu
The word "token" is a bit misleading. A token from a 0.6B small model or an embedding model is not equivalent in compute consumption or intelligence level to a token from a state-of-the-art large model capable of deep reasoning.
With this OpenClaw wave, everyone is using very SOTA open-source or closed-source models, so their token quality is very high.
👦🏻 Koji
On Bailian, roughly how many users consume 100 million tokens daily?
👨🏻💻 Wenyuan Yu
We feel like it's increasing every day, numbering in the tens of thousands. So the threshold for your "100 Million Token Club" might need to be raised.
👦🏻 Koji
Is it becoming the "Billion Token Club"?
👨🏻💻 Wenyuan Yu
I think so. Because many of our Coding Plan users are individuals who heavily consume tokens, so 100 million is no longer a very high bar.
👦🏻 Koji
Beyond token volume, what else do you care about?
👨🏻💻 Wenyuan Yu
We care about peak call volumes, how to technically smooth out peaks and fill valleys, and how to schedule effectively to fully utilize GPUs.
We also want to serve more users with better service quality, including time-to-first-token latency and generation speed.
👦🏻 Koji
Is going international a good way to keep GPUs running 24/7? For example, Chinese users during China's daytime, European users during Europe's daytime.
👨🏻💻 Wenyuan Yu
Token export is definitely a very important initiative — we're just getting started.
Alibaba Cloud is very committed to internationalization, but domestic and international business don't grow at the same pace. If you're even slightly two months behind, overseas business might only account for single-digit percentage. While it's hard to advance both in parallel, I believe the eventual middle platform will be unified for domestic and international use.
Of course, this requires overcoming many issues — geopolitics, compliance, and so on. But overall, AI going global for Chinese companies is an unstoppable, inevitable direction.
👦🏻 Koji
At Bailian, do you have a kind of "God's-eye view" — seeing which sectors and scenarios consume the most tokens? Any interesting stories to share?
👨🏻💻 Wenyuan Yu
For example, there's a beverage manufacturer that built a bot in their distributor group chat. When distributors want to restock, they simply message the bot in the group: "X cases of such-and-such drink," and the bot understands what drink it is, their past purchasing habits, and directly arranges the restock. The entire process happens through very natural language in a single group chat.
👦🏻 Koji
That does sound like a more natural way of working — no need to learn new systems, just like communicating with a real person.
👨🏻💻 Wenyuan Yu
This is certainly an inevitable future. When large models truly begin replacing the role of "people," the way many industries operate will be profoundly transformed.
👦🏻 Koji
Today all cloud providers have launched their own MaaS services, and from the outside they seem largely similar. From an insider's perspective, where does differentiation mainly lie? What might be the winning factor?
👨🏻💻 Wenyuan Yu
I'm a hands-on engineer who believes in the "hands-dirty club" philosophy. A company's infrastructure and technical execution greatly affect the final product.
Alibaba Cloud has been building infrastructure domestically for a long time, accumulating strong scale and technical depth. At the same time, we have partners like Tongyi Lab that provide us with good, non-black-box models — we polish them back-to-back together.
We also have teams like T-Head, our chip team. I've used many of their chips since the CPU era, and the development experience and efficiency have been excellent. From models to infrastructure to compute scale to self-developed hardware, we have the ability to do end-to-end polishing — that's a very unique position.
👦🏻 Koji
Specifically, when calling Qwen models on Bailian versus other platforms, what's your unique advantage?
👨🏻💻 Wenyuan Yu
I've heard many customers ask: why does their self-deployed Qwen model underperform Bailian's in terms of quality, accuracy, or speed? And this isn't limited to Qwen — it applies to open-source models too.
Because over the years we've built a very mature inference serving framework with guarantees for precision and stability. We can confidently say that for all Qwen models, the official scores published on their Model Cards can definitely be reproduced through Bailian's API.
Build Your Own or Go Cloud? Let Me Make a "Hot Take"
👦🏻 Koji
Today some larger enterprises feel that private deployment and building their own GPU infrastructure is cheaper. In your view, are there scenarios where enterprises should consider building their own?
👨🏻💻 Wenyuan Yu
Let me make a "hot take" — not representing Alibaba Cloud, but specifically Alibaba Cloud Bailian: I don't think there's any situation that requires building your own.
👦🏻 Koji
Does this also align with the trend toward increasingly specialized social division of labor? Professionals doing professional things.

👨🏻💻 Wenyuan Yu
On one hand, yes. On the other hand, people are underestimating the complexity and growth speed of this.
Customers procure GPUs for essentially three reasons:
First, cost control. I have this many GPUs, my spending for the next few months is predictable.
Second, security. My models, data, and requests are my own — not going through a third-party API feels safer.
Third, flexibility. I bought NVIDIA cards, feels like I can deploy any model, adjust my business however I want.
But my view is exactly the opposite. If these are your three motivations, then MaaS is actually the best choice.
👦🏻 Koji
Cost, security, flexibility — you believe MaaS can actually satisfy these three needs better than building your own?
👨🏻💻 Wenyuan Yu
First, cost. If you want to calculate cost per token, you need to solve several problems.
The first is inference optimization. Models and algorithms iterate quickly — it's hard for every company to maintain an infrastructure engineer to continuously optimize inference efficiency.
The second is resource utilization. Can you guarantee that your own GPUs are always running efficiently? With more and more models, finding the balance between service quality and cost is a very complex systems problem.
Second, security. As a cloud provider, trustworthiness is our bottom line — we cannot and will not look at user data. We're also promoting something called "confidential computing." In this mode, we can't even see your model files or request content. End-to-end encryption keys are in your hands alone — this is cryptography-grade protection.
Finally, flexibility. The biggest certainty today is uncertainty. You don't know what AI will need tomorrow, or how model architectures will change. For MaaS users, what they face is always a simple API.
Non-Consensus: Don't Let AI Write Too Much Code for You
👦🏻 Koji
Quick question — your undergraduate degree was in computer science. For today's juniors and seniors, would you still recommend they study computer science? If so, how do you think their learning approach should differ from yours?
👨🏻💻 Wenyuan Yu
My recommendation is yes, continue studying computer science. My advisor was an undergraduate at Peking University in the 1980s. He said his professor told him: In the future, there will only be two kinds of people in the world — those who use computers, and those who are used by computers.
👦🏻 Koji
Very accurate. Like that LatePost article, Trapped in the System.
👨🏻💻 Wenyuan Yu
Right. This was true in the 1980s, true when I was in college in the early 2000s, and remains true today. Computer science students need to understand that no matter what problems AI solves in the future, it ultimately has to manifest in the physical world — in the design and production of logic circuits and silicon chips.
AI will play an increasingly large role in this process, but we can't afford to not know how any of this happens.
As for how to use AI, I advise juniors and seniors not to use AI to write too much code for them. This may seem counterintuitive.
I saw an interview with Wen Zhang the other day about how doctors should use AI. Senior doctors using AI is definitely fine. But if an intern doctor, from their very first patient, just dumps cases into AI, they'll be unable to catch AI's mistakes.
Because they haven't accumulated experience, they lack judgment about what's good or bad, right or wrong. They can only trust AI's 99% accuracy rate, and miss that 1% of problems.
New computer science graduates must avoid becoming that 99% who overlap heavily with AI without unique skills. Instead, strive to become someone who can identify that 1% that AI can't do.
👦🏻 Koji
That's interesting. Cursor recently released data showing that six months ago, only 20-30% of users blindly accepted code completions. Now that ratio has flipped to 70-80%. This trend seems irreversible.
Can you give a specific example? Because you've handwritten so much code yourself, developing your own judgment and taste, so at some point you had a different view from an AI-generated solution.
👨🏻💻 Wenyuan Yu
This happens constantly in our daily work. When doing code review, if I see at a glance that it's AI-generated code, at the very least I get nervous. We have too many cases of deleting AI-submitted code.
AI coding for building a prototype — quality is completely fine. But for production environments, it's still a bit short. When the system truly faces pressure, problems will surface.
For production-grade code, you need to clearly understand what every single line accomplishes and that its side effects fall within acceptable bounds — that it won't cause memory leaks or consume excessive file handles. AI's depth of understanding around context and business logic is still far from reaching that level.
So I believe that for mission-critical code, we still can't fully rely on AI for now. It's an efficiency tool, but you can't assume it can do everything. That judgment matters enormously. Otherwise what you're bringing in isn't an efficiency tool — it's a burden that leaves you struggling through a mountain of shitty code.
👦🏻 Koji
Struggling through a sea of suffering?
👨🏻💻 Wenyuan Yu
Exactly. I think spec coding is a better approach — you write extremely clear requirement documents or specifications. This places extremely high demands on architects.
At last year's top storage conference, FAST, there was a paper where researchers had AI write a file system. They provided all kinds of specifications to the AI in crystal-clear form, and found that even the 32B model available at the time could write high-quality low-level code like a file system, as long as the specifications were written clearly enough.
That was hugely inspiring to me: if a person can use a somewhat formalized logic to clearly describe what they want, AI can definitely do an excellent job filling in the blanks. But I wouldn't dare say that today you can get it done well with just two or three prompts.
So I encourage everyone to experiment more with AI, but the precondition is that you yourself must first possess the capability to complete that task.
👦🏻 Koji
Some companies are quite aggressive these days, treating "increasing the proportion of AI-generated code" as a KPI. In your view, is this somewhat dangerous?
👨🏻💻 Wenyuan Yu
I think it's extremely dangerous. Anyone with even a slight understanding of the limitations of today's AI algorithms knows this is dangerous. Collaboration between humans involves so much knowledge transfer that is tacit and process-based — it can't be explained clearly in a few prompt sentences.
Many decisions, like how code should be written, may be deeply connected to a founder's style or a product's historical context — things that AI simply cannot get.
👦🏻 Koji
Don't underestimate AI's capabilities, but don't overestimate them either.
👨🏻💻 Wenyuan Yu
Absolutely don't overestimate AI's capabilities. But I do believe it's an excellent efficiency tool.
What we should be pursuing today is an engineer who, with AI assistance, can accomplish what used to require several engineers — not using one AI to replace several engineers.
👦🏻 Koji
I've also been thinking recently about the importance of "process knowledge." Three elements are needed to accomplish something: production factors, knowledge factors, and process factors. For example, when we go to IKEA to buy furniture, we acquire the production factors (parts) and knowledge factors (instruction manual), but the assembly process is still painful.
Yet a skilled craftsman shows up at your door and knocks it out in no time. He has the same factors you do, but because he's repeated this countless times, he has mastered the process knowledge.
👨🏻💻 Wenyuan Yu
Completely correct. So personally, I feel that one crucial thing for programmers is that you must build your capabilities in areas where AI cannot compete. Even if AI someday reaches 99.9% capability, you must hold onto your 0.1%.
The Most Counterintuitive Prediction: Operating System Developers Will Be the First Replaced by AI
👨🏻💻 Wenyuan Yu
You'll find that many things about AI are deeply counterintuitive. For instance, after seeing that AI could write a file system, my stronger intuition was: AI may replace first those who write the most elite code — engineers working on operating system kernels, database kernels, file systems. These domains are the easiest to be massively improved and replaced.
👦🏻 Koji
This is the opposite of what many people think?
👨🏻💻 Wenyuan Yu
Right. Frontend engineers, or roles tightly integrated with product — much of their logic isn't simple replication but requires a know-how, an understanding of how to help users actually use the product better.
👦🏻 Koji
I see. The closer something is to human interaction, the harder it is to replace. While systems engineers, because they focus on maximizing system utilization with very clear objectives, are actually easier to replace.
👨🏻💻 Wenyuan Yu
Yes, they're easier to replace. And those domains have very high-quality codebases, with very clear test cases and results — the problems are like precisely defined math problems.
AI today performs extremely well in math competitions and programming competitions precisely because these problems have sufficiently clear objectives. But "what makes a good short-video app?" — there's no clear definition for that.
👦🏻 Koji
That's an open question?
👨🏻💻 Wenyuan Yu
Very open.
👦🏻 Koji
Then do you think MaaS systems engineers face an open question or a closed one?
👨🏻💻 Wenyuan Yu
I actually think it's an open question. Because AI is changing so rapidly today — I don't know what tomorrow will look like. In this situation, underlying resources and compute availability are also changing fast.
What a systems engineer needs is no longer fixed knowledge, but the potential to respond to change.
Compute Is Like Oil, but Today's Key Isn't Just Compute
👦🏻 Koji
Today when we talk about China's AI, we can't avoid NVIDIA's chip supply issues. In your view, how significant is the impact?
👨🏻💻 Wenyuan Yu
The impact is huge, very huge. I have tremendous confidence in domestic compute — I believe it will definitely become excellent. But in today's market, compute is somewhat like oil.
The essence of the problem isn't whether China can produce oil, but whether China's daily oil needs match what it can supply. I completely believe China can achieve self-sufficiency and independent control. We have brilliant engineers and a powerful industrial base — technically, we can absolutely become world number one.

But the oil fields aren't fully exploited yet, while downstream demand is already enormous.
👦🏻 Koji
Like cars already running on the highway, but there isn't enough gas?
👨🏻💻 Wenyuan Yu
Right. If compute supply falls short, it will seriously affect China's AI development. Many countries face similar issues, though they may be bottlenecked on electricity.
Electricity is the blood of industry — this was understood 100 years ago, yet today we still face "insufficient blood supply." From my personal perspective, any compute supply that can come in is nothing but beneficial for China.
👦🏻 Koji
Among domestic chips, like T-Head, and some newly listed companies, which ones do you think are doing well?
👨🏻💻 Wenyuan Yu
T-Head is doing very well, really very well. NVIDIA is the de facto standard — it started early and has very strong hardware, software, and design teams. T-Head is the smoothest, most user-friendly domestic chip we've worked with.
👦🏻 Koji
Then what about Moore Threads and MetaX, which are hot in the capital markets?
👨🏻💻 Wenyuan Yu
I haven't personally used them. But still, today's MaaS and AI development doesn't depend on whether a particular compute product is good or not — just as whether oil is light or heavy crude, someone can always refine it.
What we face today is a total supply problem. You say my compute can grow 10x next year — why not 100x? Give me 100x the compute, and the market will definitely consume it.
Even give me 1000x, and I believe the market can use it up by training stronger models or making more AI applications cheaper.
So I always feel compute is insufficient. People may think China's compute, including Alibaba Cloud's compute, is already enough — but it's actually not, it could be more.
👦🏻 Koji
This feeling is unprecedented in the cloud computing era, right? Never before did we feel like no matter how much cloud resource you provide, it could all be used up.
👨🏻💻 Wenyuan Yu
Historically, there has never been such explosive demand for compute. Alibaba Cloud had periods of rapid growth, but this kind of "thirst" for compute like today is unprecedented.
👦🏻 Koji
Let's make a prediction: by the end of 2026, what new or even unexpected scenarios do you estimate will consume massive amounts of tokens?
👨🏻💻 Wenyuan Yu
My threshold for "unexpected" is already extremely high now — I feel like nothing AI does would surprise me anymore.
Agent is definitely one of the biggest increments this year, and AI generation definitely is too. I can't say which will be larger — each vendor's situation may differ — but I believe these two directions will definitely be this year's biggest growth points.
👦🏻 Koji
Users today can also bypass you entirely and directly call OpenAI or Qwen's API. As an intermediary platform, how thick is the value you provide?
👨🏻💻 Wenyuan Yu
First, Qwen's API is Bailing's API. The thickness of value we provide lies in who can deliver better experience, lower cost, and better model performance.
Additionally, there's capacity. We need to ensure not a single GPU sits idle for even a second, and efficiently convert that compute into tokens. Today's MaaS is essentially a compute-to-token converter — whoever converts more efficiently and whoever has more compute has the advantage.
👦🏻 Koji
On Bailing, besides Qwen, can you use other models?
👨🏻💻 Wenyuan Yu
Of course. As a cloud platform, Bailing serves customer needs. China's mainstream open-source models all have hosted deployments on Bailing.
Furthermore, for domestic model providers like MiniMax, Moonshot AI, and SiliconFlow's DeepSeek, users can also use Bailing's API key to call their model services.
👦🏻 Koji
Earlier we talked about Neocloud. Among these new cloud providers, are there any you're particularly optimistic about?
👨🏻💻 Wenyuan Yu
Neocloud is a very broad concept — everyone essentially wants to shield customers from some complexity. I'm personally not optimistic about Neoclouds that simply do resource resale, like relatively bare, low-level resale of NVIDIA compute.
I'm more optimistic about AI-native companies that deeply abstract away hardware and complexity, like MaaS providers such as Fireworks and Together. There are also interesting companies around the AI Agent ecosystem — like those doing sandbox hosting, cloud desktop browsers, and Agent observability like Datadog.
Around Agent and AI cloud-related infrastructure, very interesting companies and products will emerge.
The Future Infrastructure Stack: Water, Electricity, Gas — and Models
Koji
One last question. MaaS is still in a fiercely contested battleground today, with changes happening by the day. Have you ever thought about what kind of shift would finally settle this war, if we look to the day when the dust eventually clears?
Wenyuan
It fundamentally depends on what role AI will play in society as a whole. I believe it will become a public utility on par with water, electricity, and gas — like telecom carriers or highways, a piece of infrastructure.

But the endgame for AI probably isn't a single model. With electricity, I don't distinguish between nuclear and hydro — I plug into the socket and get 220V AC.
It will inevitably be a highly diverse, complex ecosystem, taking many forms across speed, model performance, and functionality. It probably won't be that monopolistic.
Koji
Do you think future infrastructure will evolve from "water, electricity, and gas" to "water, electricity, gas, and models"?
Wenyuan
Absolutely, without question. It will profoundly reshape our daily lives.
Koji
Great. It's been a real pleasure having you on today, Wenyuan.
We're living through a time of rapid change, and I'm very much looking forward to seeing what new developments emerge when we speak again in six months or a year.
Wenyuan
Sounds good. Thank you.
