Megatrend · Artificial Intelligence
When Nvidia's biggest customers got tired of paying the toll, they started designing their own chips
Google, Amazon, Microsoft, and Meta are the biggest buyers of AI chips from Nvidia in the world — and they're also the ones who most want to stop depending on Nvidia. Their answer is to design their own chips (Google TPU, Amazon Trainium, Microsoft Maia, Meta MTIA): custom silicon that strips out everything unnecessary and keeps only their own AI workload. Cheaper, more power-efficient — but far less flexible. This lesson looks at why these "made-to-order" chips are shaking the market Nvidia rules, and who the players behind the scenes are that help these giants build them.
01What it is (the chips cloud giants design themselves)
When we talk about AI chips, most people picture Nvidia's GPUs that everyone is fighting to buy. But there's another kind of AI chip that never goes on sale — because the companies that make it make it to use themselves. This node is about that group: custom silicon / ASIC that the cloud giants design specifically to run their own AI workloads.
The word ASIC (Application-Specific Integrated Circuit) literally means "an integrated circuit designed for a specific job." Unlike a GPU, which is a "general-purpose" chip that does many things, an ASIC strips out every function it won't use and keeps only the compute units the AI workload actually needs. The result: on the job it's designed for, it's faster, more power-efficient, and cheaper per computation — at the cost of not being able to do anything else, and being much harder to develop.
On the megatrend map, this node is a leaf under AI Compute & Accelerator Silicon, inside the big trend Artificial Intelligence. Its siblings next door are GPU & Merchant Accelerators (general-purpose chips sold to everyone — the world of Nvidia/AMD) and Inference-Optimized Silicon (chips focused on "running" models). The key dividing line: those two are chips sold on the market, while this node is chips made to order for your own use.
Merchant chip = a chip made to sell to anyone on the market (like Nvidia's GPU); every customer buys the same one · Custom ASIC / in-house silicon = a chip designed for one customer only, for their own workload, never sold — Google's TPU, Amazon's Trainium, Microsoft's Maia, and Meta's MTIA are all in this group · Hyperscaler = a company that runs cloud services at massive scale (Google, Amazon, Microsoft, Meta) — the biggest customers in the AI chip business.
The point to stress: this node is not about building factories or manufacturing chips — that's Semiconductors, where TSMC manufactures for everyone. This node is about the strategic decision of "why do the tech giants choose to design their own chips instead of buying from Nvidia" — and how that shakes the entire AI chip market.
02Why it matters — escaping Nvidia's toll
The whole reason comes down to one word: margin. In the late-2025 quarter, Nvidia's gross margin sat at around 73% — meaning for every chip priced at $100, the real cost is about $27, and the rest is Nvidia's profit. And the ones paying that "toll" are the cloud giants buying hundreds of thousands of AI chips a year. For a company pouring tens of billions of dollars of capex a year into AI chips, even a few tens of percent of cost savings is enormous money.
This is the incentive that makes them willing to spend years designing their own chips. The real numbers make it vivid: Google says its Ironwood chip (TPU gen 7) has a total cost of ownership (TCO) about 44% lower than an Nvidia GB200 server. Broadcom's custom chips (collectively called XPU) claim TCO savings of 30–50% on specialized work at large scale. And there's a famous real example — Midjourney moved its workload from Nvidia GPUs to Google TPUs and cut its monthly compute cost from $2.1 million to $700,000, roughly 65% cheaper.
The second reason is power. AI data centers eat so much electricity they're starting to hit the limits of the power grid. A custom chip designed for one job can be far more power-efficient per computation — Google's Ironwood draws about 157 watts each versus roughly 700 watts for an Nvidia B200. When you run hundreds of thousands, even millions, of chips at once, that power gap translates directly into the electricity bill and the ceiling on how far a data center can expand — connecting straight to Energy Transition & Power.
The third reason is controlling supply. In an era when Nvidia chips are scarce and you have to wait in line across years, having your own chips means not fighting others for stock — and not revealing your internal AI plans to a competitor who sells chips. For cloud giants that compete fiercely with each other, this is a big deal.
03How it works (from workload to chip in the data center)
The question many people wonder about is this — if Google or Amazon aren't "chip makers," how do they make chips themselves? The answer is that they don't do it alone, and they don't build the factory themselves. They use a "co-design" model that splits the work clearly. Let's walk through it step by step.
The heart of the mechanism is step 2 — the co-design. The cloud giants have something ordinary chip companies don't: they know exactly what their own AI workload looks like. How Google's model works, what algorithm Meta uses to recommend content — that knowledge lets them design a chip that "fits" the job. But they lack a chip-design engineering team and a library of IP (like high-speed SerDes interconnect circuits), so they rely on Broadcom or Marvell, who have all of that.
This is why Broadcom and Marvell became the "quiet winners" of this field — they don't compete with Nvidia head-on; they sell the service of "helping cloud giants build Nvidia's rivals." Once the cloud giants finish the design, they send the design files to TSMC to manufacture on 3-nanometer-class technology, then get the chips back to install in their own data centers. This cycle takes years and billions of dollars per chip generation.
04Where it sits in the ecosystem
This node sits at an interesting crossroads, because it connects both to its siblings in the same AI-chip layer and to other big trends sitting in a completely different corner.
- Rival and opposite of GPU & Merchant Accelerators: a GPU is a "buy-it-ready" chip that's flexible and has the CUDA software moat, while a custom ASIC is a "made-to-order" chip that's cheaper and more power-efficient but can't do anything else — both live together in the same data center: the GPU handles the hard, flexible work, the ASIC handles the high-volume repetitive work
- Overlaps with Inference-Optimized Silicon: almost all of the cloud giants' custom chips focus on "running" models (inference), because that work is repetitive and predictable, which suits a specialized chip — so the line between these two nodes is very thin
- Depends on Semiconductors to manufacture: the cloud giants can design but can't manufacture themselves; almost every cutting-edge custom chip is made at TSMC on 3-nanometer technology, so this node is a huge "demand" feeding the chip industry without having to compete in manufacturing itself
- The real customers are Hyperscalers / cloud providers: unlike a GPU sold to anyone, this group has only a handful of customers — and the customer is also the designer. This is a node where "buyer, designer, and user" are the same company
- Connects straight to Energy Transition & Power: since one of the main reasons to make your own chip is "saving power," this node is inseparable from the data center's energy crunch
05Where it stands now
2025–2026 is when custom chips "proved themselves" — not just experimental projects, but a real business worth tens of billions. The clear leader is Google, which has made TPUs since 2015 and is now on its 7th generation, Ironwood (TPU v7) — this generation links up to 9,216 chips per pod (many times more than Nvidia's setup, which tops out at 72 GPUs per system) and delivers 4,614 TFLOPS each, right on par with the Nvidia B200. Crucially, Anthropic, the maker of Claude, announced it would use as many as 1 million TPUs to run its own models.
The fastest-growing player behind the scenes is Broadcom. Broadcom's AI chip revenue in Q1 of fiscal 2026 hit $8.4 billion, up 106% year over year, and the company revealed a backlog of around $73 billion, while targeting AI chip revenue above $100 billion in 2027. The most talked-about deal is an agreement with OpenAI to build 10-gigawatt custom chips, which analysts value at over $100 billion.
The other cloud giants are right behind — Amazon has the training chip Trainium (gen 3 on 3-nanometer technology) and the inference chip Inferentia; Microsoft launched Maia 200 focused specifically on inference, also on 3-nanometer; and Meta has MTIA, already deployed in the hundreds of thousands to run content-recommendation systems on Facebook and Instagram — the players behind Amazon and Microsoft are mostly Marvell, which holds about 20–25% of the ASIC-design market (Broadcom holds 60–80%).
06The road ahead
The first direction is that the share of custom chips keeps rising. In 2024, custom ASICs were only about 8% of the whole AI accelerator chip market, but many firms expect that to climb to roughly 19% by 2033 — in money terms, growing from about $13 billion (2024) to over $150 billion (2030), at nearly 50% a year on average — faster than GPU growth.
But this doesn't mean GPUs lose — because the whole cake is growing faster than either side can eat it. Many firms still see GPUs holding around 70–81% of the AI accelerator market in 2030. So the clearest direction isn't "ASIC kills GPU" but a division of labor — GPUs take the hard, flexible model-training work, while custom ASICs take the massive, repetitive inference work inside the cloud giants' own house.
The second direction is competing on "how fast you ship a new generation". Cloud giants now release a new chip generation almost every year — Amazon announced Trainium4 (3× faster) for late 2026 / early 2027, and Meta has mapped MTIA out to a 500-series — a rhythm that turns Broadcom and Marvell, who sit behind everyone, into businesses with steady, predictable revenue (unlike selling chips in lots that rise and fall with the cycle).
The third direction is custom chips starting to "sell to others". The line between custom and merchant is starting to blur — there are reports that Meta is interested in using Amazon's Trainium chips, and that Google is starting to offer TPUs for outside customers to rent. If this trend grows, chips once "made for your own use" could become products that truly compete with Nvidia on the open market.
07Challenges & risks
The first risk is the CUDA moat and flexibility. Nvidia's biggest advantage isn't the chip, it's the software — CUDA, which AI developers worldwide know well and which nearly every library is tuned to run on. The cloud giants' custom chips have to build their own software tools in parallel, which is hard and takes years — and as long as the work demands "flexibility" (like new model architectures that change fast), a GPU that can do anything keeps its edge over an ASIC locked to the old job.
An ASIC requires investing in the design and the mold years ahead before it's actually used. If the AI model architecture changes in that time (say, from one type to another), a chip designed for the old job can become "obsolete" before it pays back — a risk a more flexible GPU rarely faces.
The second risk is the huge upfront cost, which only pays off for the big players. Designing one chip generation eats billions of dollars and years of time. It only pays off when there's enough volume of work to spread the cost thin — meaning this node is a game for a handful of giants with large enough data centers. Smaller companies still have to rely on GPUs on the market.
The third risk is concentration at Broadcom, Marvell, and TSMC. When nearly every custom chip from the cloud giants has to pass through just two co-designers (Broadcom + Marvell control over 80% of the market combined) and is made almost entirely at TSMC — a node born to "reduce dependence on Nvidia" has instead created a new concentrated dependence. If any one point stumbles, the whole chain shakes.
In short: this node is the story of the day the biggest customers decided "I've paid enough toll" and set out to build their own alternative — not to sell against anyone, but to control their own cost and fate in an era when everything runs on AI chips. Understand this node, and you understand why even the most powerful companies in the world still don't want to stake their lives on a single supplier.