Behind the $5 Billion: A New Round of Reshuffling in AI Infrastructure | Vital Views
Counselor Vitality

In the ongoing evolution of global AI technology, semiconductors — and AI chips in particular — serve as a critical lever for this deep restructuring. They underpin the foundational architecture of next-generation intelligent systems and shape how the entire tech value chain collaborates and innovates.
NVIDIA's recent $5 billion investment in Intel may appear, on the surface, as a capital partnership that transcends old competitive narratives. In reality, it looks more like a preemptive move to secure AI infrastructure, kicking off a new wave of system-level restructuring worldwide — from hardware packaging to cloud service architecture to data center deployment logic, everything is quietly shifting.
We've chosen to republish this translated article in the hope of catching, together with you, the signals beneath these surface-level clues: When AI is no longer merely an evolution of "technology" but becomes a migration of the entire tech paradigm, what angles can help us understand the rhythm of present change and capture new trends and possibilities?
Enjoy
In the latest episode of the a16z (Andreessen Horowitz) video podcast, SemiAnalysis chief analyst Dylan Patel, a16z partner Sarah Wang, and a16z partner Guido Appenzeller — former CTO of Intel's Data Center and AI Group — jointly explored the significance of this deal, NVIDIA's moat, Huawei's rise, and the future trajectory of AI infrastructure, revealing a new era driven by the interplay of technological innovation, geopolitics, and trillions in capital.

Core Takeaways
-
NVIDIA's Intel investment reshapes the landscape: This is a brilliant strategic move for NVIDIA, one that promises substantial returns while piling enormous pressure on rivals like AMD and ARM, reshaping competitive dynamics across PC and data center markets.
-
Huawei is NVIDIA's most formidable challenger: Despite US sanctions, Huawei is rapidly closing the gap in AI chip design and manufacturing. Its formidable execution capabilities and domestic market backing make it a competitor that NVIDIA cannot afford to ignore globally, especially outside the US market.
-
AI compute demand drives a trillion-dollar market: Hyperscale data center construction is proceeding at an unprecedented pace. Oracle, with its nimble strategy and massive balance sheet, has successfully landed major customers like OpenAI, emerging as a big winner in the AI cloud market.
-
NVIDIA's success stems from high-risk bets: Jensen Huang, with his extraordinary intuition and willingness to wager the entire company, has led NVIDIA to repeatedly catch market tailwinds and build a formidable technological and market moat.
-
New hardware brings new challenges: While next-generation GPUs like the GB200 deliver massive performance leaps, their high total cost of ownership, complex deployment requirements, and reliability issues pose fresh infrastructure challenges for users.

When long-time rivals shake hands, the market always senses something unusual. NVIDIA's $5 billion investment in Intel immediately drew widespread industry attention upon announcement. Patel noted that this investment is a masterstroke for NVIDIA — Intel's stock jumped 30% on the news alone, generating substantial potential returns for NVIDIA.
The deeper significance lies in the fact that this collaboration extends beyond financial investment to encompass joint development of custom data center and PC products, which Patel described as everything coming full circle, with Intel somewhat groveling to NVIDIA. He recalled the litigation between Intel and NVIDIA over alleged anti-competitive behavior in chipsets; now Intel is manufacturing chiplets to be packaged with NVIDIA GPUs for PC products — an undeniable reversal of market dynamics. Patel believes these deeply integrated x86 laptops could become "the best products on the market."
Former Intel executive Guido Appenzeller put it this way: "If your two biggest mortal enemies suddenly team up, that's the worst news you can get." He sees short-term benefits for customers and consumers, particularly in the laptop market. Yet the impact on competitors is seismic. Appenzeller didn't mince words: both AMD and ARM face serious challenges. AMD's graphics cards are decent, but its software ecosystem is weak, and the NVIDIA-Intel alliance will further squeeze its room to survive. For ARM, its core selling point was its non-cooperation with Intel — but now NVIDIA may use this opening to enter Intel's technology domain and become a far more dangerous CPU competitor. Appenzeller concluded: "I didn't expect the deck to be reshuffled like this. I think it's an amazing development."

On the other side of global chip competition, despite brutal US sanctions, Huawei's AI chip business is rising against the headwinds to become NVIDIA's most serious challenger in global markets, particularly outside the United States. Patel recalled Huawei's glory days in 2020 when it launched the Ascend chips — at the time, they were the first company to ship 7nm AI chips, with a technical gap relative to NVIDIA that was virtually nonexistent.
Sanctions, however, stripped Huawei of TSMC's foundry services, forcing it to turn to domestic SMIC for production while sourcing memory supplies through various overseas channels. Patel revealed that Huawei secured orders for nearly 3 million chips from TSMC through a complex web of entities, worth roughly $500 million in total. Though these channels were eventually cut off, Huawei's resilience cannot be underestimated.
With NVIDIA's China-market AI chips like the H20 banned in 2025, the Chinese government is actively pushing for domestic substitution. Patel pointed out that Chinese companies including Huawei are rapidly catching up in logic chips (replacing TSMC) and memory chips (replacing SK Hynix and Samsung). While gaps remain, he believes China can manufacture substantial volumes of 7nm AI chips — and may even produce 5nm chips using existing equipment.
More surprisingly, Huawei has also broken through in memory, announcing plans to adopt custom HBM (high-bandwidth memory). Patel sees this as proof that Huawei is rapidly catching up to NVIDIA and AMD's memory technology roadmaps. While capacity and yield remain critical bottlenecks, Patel emphasized: "The question is always, can we manufacture it? What Jensen Huang would say is, are you betting that China can't manufacture it? It's a matter of when, not if."

Against this turbulent semiconductor backdrop, the explosive growth in AI compute demand is driving an unprecedented trillion-dollar market. Patel noted that Wall Street's total capex forecast for all hyperscalers (Microsoft, CoreWeave, Amazon, Google, Oracle, Meta) next year sits around $360 billion, while his research model points closer to $450–500 billion — with the overwhelming majority flowing to NVIDIA.
In this compute arms race, Oracle has distinguished itself through a distinctive strategy.
Patel considers Oracle's agreement with OpenAI, worth over $300 billion, "the most unprecedented thing in the history of stocks and public companies." He explained Oracle's winning formula: possessing the industry's largest balance sheet while refusing to lock itself into any particular hardware or networking technology, giving it exceptional flexibility in evaluating and snapping up data center capacity.
By tracking every data center build globally — construction progress, power supply, chip deployment, and customer agreements — Patel has precisely forecasted Oracle's revenue growth. He found that Oracle is aggressively contracting and deploying massive new data centers, facilities that will provide critical AI compute for major customers like OpenAI and ByteDance in the coming years.
By contrast, Amazon AWS once found itself in "crisis" in the AI cloud space. Patel wrote in Q1 2023 that AWS's historical strength in horizontally scaled computing had become a liability in the vertically scaled AI infrastructure era, with its internal AI chip team focused on cost optimization rather than performance maximization. Though AWS revenue growth temporarily slowed, Patel predicted it would reaccelerate because Amazon still commands the largest pool of idle data center capacity, which will convert to AI revenue over the next year.
Yet next-generation AI hardware deployment brings its own challenges. While Amazon has deep experience with high-density data center technologies like liquid cooling, Guido Appenzeller raised a critical concern: "Now I might actually need a big river next door for cooling. In many regions, I simply can't get enough water. And very likely, in the same region, power becomes an issue too." Patel acknowledged efficiency concerns but emphasized that Amazon is moving fast to invest in and convert this capacity.
On the evolving relationship between Microsoft and OpenAI, Patel noted that while Microsoft was once OpenAI's exclusive compute provider, that relationship is being reshaped as OpenAI seeks diversification and strikes massive deals with Oracle. The core driver behind this is balance sheet capacity: OpenAI needs a partner with sufficient capital to shoulder enormous GPU investments, and Oracle happens to fit that bill.

NVIDIA's dominance in AI chips wasn't built overnight — it's the result of a series of "crazy," high-risk bets by founder Jensen Huang. Dylan Patel likened Huang to a "Warren Buffett effect for semiconductors" and dissected the historical strategy behind NVIDIA's formidable moat.
Patel revealed that NVIDIA experienced multiple failures in its early days, but Huang was crazy enough to bet the entire company. For the Xbox project, he ordered the required number of chips before Microsoft even placed the order — a YOLO (You Only Live Once) mentality that has threaded through NVIDIA's entire history.
During the crypto bubble, when everyone questioned whether GPU demand was real, NVIDIA successfully convinced supply chain partners that this wasn't just cryptocurrency but "durable, real demand" spanning gaming, data centers, and professional visualization. They induced suppliers to ramp production, and when the bubble burst, NVIDIA simply wrote off one quarter's inventory while competitors (like AMD) dithered and missed their window.
Another hallmark of Huang is his extraordinary intuition and foresight. Patel shared an anecdote about Huang and CFO Colette: Huang once said, "I hate spreadsheets. I don't look at them. I just know." This intuition allowed him to repeatedly anticipate market demand ahead of time, even issuing capacity forecasts higher than customers' own internal plans, and committing to non-cancellable, non-returnable orders with confidence.
More impressive still is NVIDIA's execution in chip design and manufacturing. Appenzeller noted that at Intel, they deeply envied NVIDIA because NVIDIA consistently achieved first-pass success — typically submitting only an "A0" version, while other companies (like Intel) might need as many as 15 revisions ("E2" versions), dramatically shortening time-to-market. This operational efficiency enables NVIDIA to respond rapidly to market shifts and deliver innovative products on schedule — for instance, boldly adding tensor cores to the Volta chip just months before production, thereby establishing its leadership in AI.
NVIDIA's moat isn't merely technology and market share; it's Huang's distinctive leadership philosophy of playing only to win the next game. Patel quoted Huang: "The point of playing the game is to win. The point of winning — or the reason you win — is so that you can play again." This DNA of relentlessly pursuing the next technological breakthrough rather than defending existing achievements is NVIDIA's secret to longevity.

As AI compute demand surges, global data center construction is proceeding at unprecedented speed and scale, with the industry officially entering the "gigawatt era" (1 gigawatt equals 1 billion watts). Dylan Patel observed that training GPT-3 once required 20,000 H100 GPUs and seemed impressive; now the world has at least ten clusters of 100,000 GPUs each, and analysts have grown "bored" with megawatt-scale data centers — only gigawatt-level builds excite them.
In this epic construction frenzy, Elon Musk's X AI and Colossus project stands out as a striking case study. Patel detailed Musk's astonishing feat in Memphis, Tennessee: going from purchasing a factory to deploying 100,000 GPUs in just six months, building a liquid-cooled data center. He employed a series of "crazy" innovations — placing generators directly outside the factory, digging for natural gas pipelines to secure power, deploying mobile substations — to obtain energy and cooling at maximum speed.
Musk's construction speed and first-principles thinking even allowed him to challenge local government regulations. Patel mentioned that after facing protests in Memphis, Musk simply shifted some infrastructure across the border to Mississippi, exploiting regulatory differences between states to accelerate construction. This unconventional execution is key to driving gigawatt-era data center builds.
Yet such massive AI infrastructure investment also raises new questions: NVIDIA, as the largest cash-flow company, faces an open question about capital deployment strategy given future cash surpluses running into the hundreds of billions. Regulatory scrutiny has prevented Huang from executing large acquisitions (like ARM), so where does all this capital go?
Patel believes investing in data centers and power infrastructure to resolve compute growth bottlenecks may be more strategically meaningful than investing in cloud services themselves. But he also acknowledges that Huang is actively deploying capital into new cloud and model training companies — though these are small bets that can't fully absorb NVIDIA's massive cash flow. This dilemma stands in sharp contrast to Apple under Tim Cook, where a lack of vision led to massive share buybacks and innovation stagnation.

While next-generation AI hardware delivers leapfrog performance gains, the steep costs and complex challenges of deployment are creating new headaches for users and cloud providers alike. Patel analyzed the TCO (total cost of ownership) of NVIDIA's GB200, which integrates GPU and CPU on the same chip, noting it's roughly 1.6x that of the H100. While the GB200 can deliver 6–7x performance gains on specific workloads like deep exploration inference or reinforcement learning, making the investment worthwhile, for more general-purpose tasks the improvement may be only around 2x, making returns far less compelling.
The biggest deployment challenge for the GB200 is reliability. Patel revealed that as a new GPU, the GB200 still faces reliability issues, and its "blast radius" upon failure far exceeds previous generations. In a GB200 rack with 72 GPUs, if one GPU fails, the entire rack may need to go offline for repair. This differs from the 8-GPU H100/H200 boxes, where a failure only requires replacing a single server.
To address this, users and cloud providers have had to design complex infrastructure management strategies — for instance, running high-priority workloads on 64 GPUs while using the remaining 8 for low-priority tasks, reducing the impact of failures on core business. Cloud providers have adjusted their service level agreements (SLAs) accordingly, offering 99% availability on 64 GPUs rather than 72. This ultimately forces customers to develop the capability and intelligence to handle such unreliability.
Patel also elaborated on the breakdown of inference workloads: prefill and decode. These two operations have fundamentally different requirements — the former demands massive floating-point compute, while the latter is sensitive to latency and tokens per second (TPS). Top labs like OpenAI and Anthropic, along with numerous cloud providers, have begun separating prefill and decode workloads to run in parallel on different GPU clusters, improving efficiency and user experience.
NVIDIA has followed this trend with dedicated prefill chips like the Rubin prefill card (CPX). By stripping away HBM — which accounts for over half of GPU cost — CPX offers a cheaper, more efficient prefill chip that effectively supports deployment of long-context AI models. This not only reduces costs but improves overall inference efficiency.

In the rapidly evolving frontier of AI, RL (reinforcement learning) environments — key technology for training AI agents — are drawing massive attention across Silicon Valley. The article notes that RL environments aim to train AI agents through multi-step tasks simulating real software applications, which is essential for developing more powerful general-purpose AI agents.
Leading AI labs like Anthropic are pouring heavily into building their own RL environments, while startups like Mercor and Mechanize, along with traditional data labeling giant Scale AI, are actively positioning to become "the Scale AI of RL environments."
Yet the future of RL environments is no sure thing. Patel stressed that "people underestimate how hard it is to scale [RL] environments." Some industry experts point out that RL environments are prone to "reward hacking" — AI models gaming the system to obtain rewards without genuinely completing tasks. OpenAI executives remain cautious about RL environment startups, and while AI researcher Andrej Karpathy is bullish on the potential of "environments and agent interaction," he's reserved about how much progress "reinforcement learning itself" can deliver.
This suggests that while RL environments are seen as critical to AI agent breakthroughs, their technical challenges, scaling costs, and effectiveness remain to be proven over time.

As Patel put it, predicting beyond five years is nearly impossible because every situation is an entirely new game. In this Wild West woven from technology, capital, and geopolitics, Huang's philosophy that "the point of playing the game is to win, and the point of winning is so that you can play again" may be the most fitting epigraph for this high-stakes transformation.
This article is republished from the WeChat account Web3 Sky City; for the video and Chinese translation, please visit the original source.
(Disclaimer: Views are for reference only and do not represent institutional positions.)





