OpenAI Designed a Chip in Just 9 Months — The AI Chip Game Enters Its Second Half | 5Y Research
In June 2026, OpenAI and Broadcom jointly released an inference chip called Jalapeño.
Michael | Investor, 5Y Capital
In June 2026, OpenAI and Broadcom announced a推理 chip called Jalapeño.
Its performance specs were barely disclosed. But two things matter far more:
First, this chip was defined entirely from scratch by OpenAI — based on its own model architecture, future roadmap, kernel and serving stack — with Broadcom handling the silicon implementation. OpenAI no longer wanted to buy generic compute off the shelf; it wanted a chip tailored to its own body.
Second, this advanced-process mega-die, pushing the reticle limit, went from initial design to tape-out in just nine months. By the companies' account, that's the fastest development record ever for a high-performance chip. One of the accelerants: OpenAI's own models.
The first half of the AI chip era has ended, and the biggest winner is unquestionably NVIDIA. The plot was simple: to reach AGI as soon as possible, everyone was buying the best AI training machines on Earth at any cost — and NVIDIA was the only option.
Jalapeño is likely the opening whistle of the second half. It points to two things: one, as inference overtakes training, starting with OpenAI, every company that wants a seat at the AI table, every model, every application, will want the chip that fits its needs best at the best price — AI chips will no longer be a single species of GPGPU, but will see a Cambrian explosion of forms.
Two, when AI begins to dominate chip design, the cost and cycle wall that has kept most companies out of custom silicon may finally be toppled. AI itself is about to become the new creation machine.
The second half will come faster and harder than the first.
The Cambrian Explosion of AI Chips
Complex, varied tasks call for a general platform that can cover all possibilities. Targeted tasks with clear boundaries demand specialized systems that push cost-performance to the extreme.
For AI, training is the former: heavy, volatile, rapidly iterating — like building aircraft carriers and fighter jets. Only the rarest players can afford it. What's needed is generality, stability, scalability, maxed-out performance; cost is a secondary concern. Inference is the latter: once the model and scenario are fixed, the task becomes narrow and deterministic — like building family cars. Good enough is the goal; cheaper is better.
Unlike the CV era, in the LLM era inference's share of total AI compute was roughly one-third in 2023, crossed half in 2025, and will reach two-thirds in 2026 — and that's before agents deploy at scale. Jensen Huang put it more bluntly: because of reasoning and agents, AI's compute needs are "easily 100× what we expected a year ago." Once agents start writing code, doing research, and handling tasks around the clock, inference demand pulling ahead of training by more than an order of magnitude is only a matter of time.
Any inference scenario you consider too niche or too small today could spawn one or two companies within a year or two, with demand large enough to justify a custom chip the way Google once did for its own workloads. And customization can go as extreme as Taalas: baking entire model weights directly into transistors. Extreme customization yields extreme results — this chip has no HBM, and on a hardened Llama 3.1 8B, claims roughly 17,000 tokens per second per user, about two orders of magnitude faster than a GPU.
How many kinds of AI inference chips do we need, and of what sort? This question resembles the eve of cars replacing horse-drawn carriages — people then also asked: how many car brands, how many models does the world need?
Such questions never have standard answers. Answers are voted up by countless specific needs: every model, every application, every hardware form factor is an independent buyer — like cars, each with its own route, payload, speed, energy consumption, budget, and road conditions. Buyers don't pick the most expensive or most powerful; they pick what fits.
The endgame for AI inference chips is necessarily highly fragmented — upstream applications demand fundamentally different things, and those demands will only diverge further; yet each category is large enough to justify a dedicated chip. All called "inference": chatbots demanding sub-second response, reasoning models thinking through tens of thousands of tokens, agents making dozens of tool calls, long-context tasks swallowing entire code repositories, hyper-realistic video generation, trillion-parameter MoEs, VLA/world models packed into robot bodies — bottlenecks land in completely different places: bandwidth, compute, memory, interconnect, power, each its own constraint.
Every row in this table already has corresponding chips or systems today, and they're looking less and less alike. Even NVIDIA is bifurcating its own product lines. And this fragmentation will keep advancing, because the core driver is the rapidly expanding inference market. Dylan Patel of SemiAnalysis says AI inference will be a market "far larger than oil"; the larger the market, the more niches for segmentation — every demand category, even every individual model, is large enough to sustain a dedicated chip or even a product family.
NVIDIA today resembles Ford a century ago — with a single, ruthlessly standardized Model T, it once captured nearly 60% of the US auto market, and expanded car ownership by two orders of magnitude; yet the very prosperity it cultivated ultimately attracted rivals who understood segmented needs better, squeezing it from "the only choice" to "one of many choices." The same script is likely to replay in chips: a market grown large enough no longer fits into a single product form.
The Next TSMC: Birth of the Creation Machine
Soon we'll face the next question: if the future needs several orders of magnitude more chip varieties, who will design them? Who can afford the hefty upfront customization costs, and who can wait out the lengthy design and tape-out cycles?
A simple economic formula makes this clear:
AC(Q) = F / Q + VC
AC(Q) is the cost amortized per chip — only when it's low enough does customizing for one scenario, one customer, one model make sense; F is the fixed cost of design, IP, and tape-out for one chip; Q is the market volume that chip can cover; VC is the marginal cost of manufacturing each unit.
How high is F? Today, a relatively simple 5nm chip reusing mature IP costs $100–200 million just to design; more complex ones run to $500 million or more. Even custom chips on mature nodes like 12nm or 28nm cost tens of millions in design alone.
Equally lethal as design cost is time — from definition to tape-out, a chip typically takes 18 months to three years; meanwhile models iterate on a monthly cadence, hardware's rhythm simply cannot keep up.
Most companies can't afford this money or wait this cycle. Yet as the AI inference market further expands, the motivation to find more cost-effective hardware for different workloads only strengthens — vast demand and innovation remain stifled at the bud.
How real is the demand this wall blocks? Look at robotics. A humanoid robot contains one to two thousand semiconductors — nearly identical joint-control MCUs alone number twenty to seventy, all running on mature processes with mask sets costing roughly $100,000. These should be the easiest chips to customize; yet every company's actuator architecture differs, fragmenting demand so that no single design reaches the break-even line of several hundred thousand units — so the entire industry still assembles from catalog parts. Everyone fixates on the 3–5nm "expensive brain"; equally stuck are the dozens of important small chips on each robot that no one is willing to design and tape out.
This scene strongly resembles the semiconductor industry in 1987. The last company to make a disruptive innovation that transformed the entire semiconductor industry was TSMC.
Back then, the industry was still IDM territory — Intel, TI, and other giants that both designed and manufactured captured the first wave of surging chip demand; yet the astronomical cost of a fab shut out legions of companies with design ideas, who could only beg for IDMs' idle capacity. TSMC, the first third-party fab that didn't do its own designs, didn't invent any new technology per se — it simply platformized wafer manufacturing capability into infrastructure that everyone in the industry could access relatively cheaply and quickly, unlocking orders of magnitude more design innovation than the IDM era. The rest is familiar history: the fab + fabless division of labor birthed today's three largest semiconductor companies by market cap — NVIDIA, TSMC, and Broadcom.
Today's situation closely parallels 1987: AI, embodied intelligence, and massive new demands are pushing semiconductor needs to new heights; model companies, application companies, hardware companies all hold fistfuls of custom chip ideas and concepts, yet everyone is blocked by the towering "F." Everyone envies Google's partnership with Broadcom, but few can afford that design services bill or survive TPU's decade of trial and error. The market's call for the next TSMC has already emerged — it should be a Broadcom that is 10× faster and 10× cheaper, becoming the central node of the next semiconductor industry division of labor after fab + fabless.
And large models themselves may be building that "10× Broadcom" right now.
When Silicon Runs on Software's Clock?
Around 2020, Google began exploring reinforcement learning for chip layout, work that later became AlphaChip and has been used across several TPU generations. OpenAI's collaboration with Broadcom explicitly noted "using OpenAI's models to accelerate portions of chip design and optimization," resulting in a from-scratch chip design in just nine months. NVIDIA Chief Scientist Bill Dally gave a figure at this year's GTC — migrating a standard cell library to a new process used to take eight engineers ten months; now one GPU runs overnight; circuits explored through their reinforcement learning are "so weird no human engineer would draw them that way," yet beat manual designs by 20–30% in area and power.
This opportunity isn't lost on smaller players either. Ricursive, founded by two AlphaChip leads, raised $335 million at a $4 billion valuation within four months of founding; ChipAgents, doing agentic chip design and verification, has raised $74 million with products in roughly 50 semiconductor companies; Cognichip, building foundational models for chips, raised $93 million with Intel CEO Lip-Bu Tan joining its board; and Agentrys, founded by NVIDIA's chip design AI lead Mark Ren, among others.
From first principles, chip design is inherently the perfect proving ground for LLM + reinforcement learning. AI iterates fastest and surpasses humans earliest in closed worlds with clear rules, instant feedback, and unlimited practice — first Go, then code and math.
Chip design may be next: the EDA toolchain itself is a sufficiently capable world simulator — RTL can be simulated, logic correctness can be formally verified, timing, power, and area can yield reliable numerical feedback without tape-out. Models need not wait for human experts to feed large amounts of high-quality data; instead they can self-explore at high speed and high throughput in simulators, as AlphaGo did through self-play. Compared to Google's era, we now also have LLMs with far stronger reasoning capabilities and world knowledge.
As long as environmental rules are clear and feedback is fast, LLMs can already take on increasingly many design steps; with rapid model capability iteration, full front-to-back agentization is not far-fetched.
Models and chips are entering a mutually reinforcing, mutually accelerating positive loop. The stronger the model, the stronger its chip design capability; it designs more efficient chips, which in turn make training and inference faster and cheaper, making models stronger faster; stronger models spawn more new scenarios, calling forth more specialized chips — each cycle turns faster than the last.

The acceleration button on this loop has already been pressed. Building chips for AI, and building chips with AI, connect head to tail. AI is for the first time beginning to design the physical substrate of its own existence.
We may be standing on the eve of the next semiconductor industry division of labor: from fabless + fab, toward designless + AI design house + fab. End model companies and application companies no longer need to understand chips — they simply submit their requirements and workload characteristics, and retrieve a custom-tailored chip as quickly and cheaply as calling cloud computing today; AI design platforms amortize customization costs across hundreds or thousands of designs, transforming chips from a hundred-million-dollar heavy-asset gamble, a game for giants only, into an ordinary order that any company can place at any time.

For sixty years, silicon's metronome was Moore's Law — two years per tick. Everyone planned around hardware's slowness and expense — product roadmaps, competitive moats, investment cycles, all built on "slow" and "expensive." When AI takes over design, chip iteration hangs on software's clock for the first time. Nine months is merely the first data point on this new curve.
There may come a day when building a chip is as ordinary as shipping a software version — and then, what will be scarce are the people who can think clearly about "what chip should be built, and why."
Welcome to the second half.


