Key points
- Training a model is a one-time cost. Inference, running it for users, is a recurring one. In 2026, inference spending passed training for the first time, according to Gartner.
- That shift rewards cheaper, power-efficient custom chips and memory over raw GPUs, which is why the big cloud companies are designing their own silicon.
- The map: Nvidia still dominates, Broadcom and Marvell design the custom chips for Google, Amazon, Meta and Microsoft, AMD is the merchant challenger, and Micron's memory rides all of it.
- Run any name here through our stock score tool to see what its filings actually say.
AI chips generate costs in two distinct stages. Training is the enormous upfront job of creating a model. Inference begins whenever a user asks that model to do something, and the meter keeps running for as long as the service is used.
Training absorbed most of the spending during the early AI boom. In 2026, the balance shifted. Gartner expects roughly $42 billion to be spent on AI-optimized cloud infrastructure this year. Inference accounts for $23.3 billion of that amount, or 55%, and its share is projected to reach 59% in 2027. By that measure, 2026 is the first year in which running models costs more than building them.
Reasoning models add another layer of demand. Working through a problem step by step produces much more text for each question, and every additional word consumes compute that someone must pay for. The inference workload is therefore increasing in two ways: more people are using AI, and each request can require more work.
Nvidia (NVDA) runs gross margins above 70%, so its biggest customers have a strong reason to design their own chips and skip that markup on the workloads they run over and over. A custom chip built for one job, an ASIC, can beat a general GPU on cost and power for that job. It can't do everything a GPU does, which is the catch.
Who makes what
| Company | Position in the inference market | Current signal |
|---|---|---|
| Nvidia (NVDA) | The standard merchant GPU and still the default for training and inference | FY2026 data-center revenue $193.7B, up 68% |
| AMD (AMD) | The merchant challenger; Instinct MI350 and MI400 accelerators | MI350 launched in 2025, MI400 due in 2026 |
| Broadcom (AVGO) | Designs custom AI processors for major cloud providers | FY2026 AI revenue about $56B; sees over $100B by FY2027 |
| Marvell (MRVL) | Co-designs custom AI silicon for major cloud providers | FY2026 revenue $8.2B, up 42%; data center about 75% of it |
| Micron (MU) | Supplies the high-bandwidth memory required by every accelerator | HBM market seen near $54.6B in 2026 (BofA estimate) |
The cloud companies are not limiting themselves to purchases from chip vendors. Google developed the TPU, Amazon has Trainium and Inferentia, Microsoft built Maia, and Meta created MTIA. Broadcom or Marvell co-designed most of those processors, turning the two suppliers into the largest custom-chip story behind Nvidia. Broadcom has also announced a chip co-design partnership with OpenAI, with products expected in 2026 and 2027.
The pressure is on Nvidia's margin, not necessarily its revenue
It's easy to read the custom-chip wave as the end of Nvidia. That's not what the numbers say. Analysts estimate Nvidia still holds most of the inference market and the vast majority of training, and its absolute revenue can keep growing even as its share slips, because the whole market is expanding faster than any rival is taking share. The ASIC story is a threat to Nvidia's fat margin, not obviously to its top line. It's a clearer win for Broadcom, Marvell and the memory makers, who get paid no matter whose chip wins.
This is also where the money actually shows up today. Nvidia, Broadcom and the memory makers get paid now, in cash, for hardware. The software companies still trying to charge for AI features have to prove customers will pay. That gap between who bills today and who bills someday is worth keeping in mind whenever the AI trade is called a bubble.
What could go wrong
Nvidia's strongest defense is software. CUDA is already familiar to AI developers, so custom chips must overcome an established working habit as well as compete on hardware. ASICs are deliberately narrow, and a change in model architecture can leave a purpose-built design stranded. Custom-silicon revenue is also concentrated among a handful of enormous customers. Finally, the investment case assumes inference demand will reach the volumes in current forecasts.
This is not a recommendation to buy or sell any security. Our stock score tool grades each company's filings in about ten seconds, while our HBM explainer examines the memory side of the market.
Sources
- Related coverage: What is HBM, and why AI memory stocks matter
- Related coverage: Inside the AI hardware stack, layer by layer
- Gartner, forecast for AI-optimized IaaS spending (August 2026)
- Nvidia, fourth-quarter fiscal 2026 results; Broadcom fiscal 2026 AI revenue commentary
- Company figures from each firm's most recent report, via SEC EDGAR
This is general market commentary and opinion, not investment advice. Markets can go down as well as up, and you can lose money. Always do your own research and consider speaking with a licensed financial professional before making any investment decision.


