For the first time, more money is going into running AI than building it. What that means for Nvidia (NVDA), Broadcom (AVGO), Marvell (MRVL) and AMD

Custom AI accelerator chips and GPUs running inference in a data center

AI-generated illustration of the chips that run AI inference in a data center

Key points

  • Training a model is a one-time cost. Inference, running it for users, is a recurring one. In 2026, inference spending passed training for the first time, according to Gartner.
  • That shift rewards cheaper, power-efficient custom chips and memory over raw GPUs, which is why the big cloud companies are designing their own silicon.
  • The map: Nvidia still dominates, Broadcom and Marvell design the custom chips for Google, Amazon, Meta and Microsoft, AMD is the merchant challenger, and Micron's memory rides all of it.
  • Run any name here through our stock score tool to see what its filings actually say.

AI chips generate costs in two distinct stages. Training is the enormous upfront job of creating a model. Inference begins whenever a user asks that model to do something, and the meter keeps running for as long as the service is used.

Training absorbed most of the spending during the early AI boom. In 2026, the balance shifted. Gartner expects roughly $42 billion to be spent on AI-optimized cloud infrastructure this year. Inference accounts for $23.3 billion of that amount, or 55%, and its share is projected to reach 59% in 2027. By that measure, 2026 is the first year in which running models costs more than building them.

Reasoning models add another layer of demand. Working through a problem step by step produces much more text for each question, and every additional word consumes compute that someone must pay for. The inference workload is therefore increasing in two ways: more people are using AI, and each request can require more work.

Nvidia (NVDA) runs gross margins above 70%, so its biggest customers have a strong reason to design their own chips and skip that markup on the workloads they run over and over. A custom chip built for one job, an ASIC, can beat a general GPU on cost and power for that job. It can't do everything a GPU does, which is the catch.

Who makes what

CompanyPosition in the inference marketCurrent signal
Nvidia (NVDA)The standard merchant GPU and still the default for training and inferenceFY2026 data-center revenue $193.7B, up 68%
AMD (AMD)The merchant challenger; Instinct MI350 and MI400 acceleratorsMI350 launched in 2025, MI400 due in 2026
Broadcom (AVGO)Designs custom AI processors for major cloud providersFY2026 AI revenue about $56B; sees over $100B by FY2027
Marvell (MRVL)Co-designs custom AI silicon for major cloud providersFY2026 revenue $8.2B, up 42%; data center about 75% of it
Micron (MU)Supplies the high-bandwidth memory required by every acceleratorHBM market seen near $54.6B in 2026 (BofA estimate)

The cloud companies are not limiting themselves to purchases from chip vendors. Google developed the TPU, Amazon has Trainium and Inferentia, Microsoft built Maia, and Meta created MTIA. Broadcom or Marvell co-designed most of those processors, turning the two suppliers into the largest custom-chip story behind Nvidia. Broadcom has also announced a chip co-design partnership with OpenAI, with products expected in 2026 and 2027.

The pressure is on Nvidia's margin, not necessarily its revenue

It's easy to read the custom-chip wave as the end of Nvidia. That's not what the numbers say. Analysts estimate Nvidia still holds most of the inference market and the vast majority of training, and its absolute revenue can keep growing even as its share slips, because the whole market is expanding faster than any rival is taking share. The ASIC story is a threat to Nvidia's fat margin, not obviously to its top line. It's a clearer win for Broadcom, Marvell and the memory makers, who get paid no matter whose chip wins.

This is also where the money actually shows up today. Nvidia, Broadcom and the memory makers get paid now, in cash, for hardware. The software companies still trying to charge for AI features have to prove customers will pay. That gap between who bills today and who bills someday is worth keeping in mind whenever the AI trade is called a bubble.

What could go wrong

Nvidia's strongest defense is software. CUDA is already familiar to AI developers, so custom chips must overcome an established working habit as well as compete on hardware. ASICs are deliberately narrow, and a change in model architecture can leave a purpose-built design stranded. Custom-silicon revenue is also concentrated among a handful of enormous customers. Finally, the investment case assumes inference demand will reach the volumes in current forecasts.

This is not a recommendation to buy or sell any security. Our stock score tool grades each company's filings in about ten seconds, while our HBM explainer examines the memory side of the market.

Sources

This is general market commentary and opinion, not investment advice. Markets can go down as well as up, and you can lose money. Always do your own research and consider speaking with a licensed financial professional before making any investment decision.

Frequently asked questions

What is the difference between AI training and inference?

Training is the one-time job of building a model by feeding it huge amounts of data. Inference is running the finished model to answer questions, which happens every time someone uses it. Training is a burst of spending; inference is a recurring cost that grows with usage. In 2026, Gartner estimates inference spending passed training for the first time. This is general market commentary and not investment advice.

Why are Google, Amazon and Meta designing their own AI chips?

To cut cost and power. Nvidia runs gross margins above 70%, so its largest customers save money by designing custom chips, called ASICs, for the workloads they run constantly. Google has its TPU, Amazon has Trainium and Inferentia, Microsoft has Maia and Meta has MTIA. Most are co-designed with Broadcom (AVGO) or Marvell (MRVL).

Will custom chips (ASICs) replace Nvidia?

Not soon, and maybe not in the way the headline suggests. Custom chips threaten Nvidia's high margins more than its revenue, because the overall market is growing faster than any rival is taking share, and analysts still put Nvidia at most of inference and the bulk of training. Nvidia's CUDA software is also a moat that custom chips have to overcome. This is general market commentary and not investment advice.

Which stocks benefit from AI inference growth?

Nvidia (NVDA) and AMD (AMD) on the merchant GPU side, Broadcom (AVGO) and Marvell (MRVL) on the custom-chip side, and Micron (MU) in memory, which every accelerator needs. Broadcom and Marvell are notable because they get paid to design chips no matter which cloud giant wins. This is general market commentary and not investment advice.

What is an AI ASIC?

ASIC stands for application-specific integrated circuit, a chip built for one job rather than general use. In AI it means a custom accelerator designed for a specific workload, which can beat a general-purpose GPU on cost and power for that task but cannot do everything a GPU does.

More on AMD and AVGO

Dennis Singleton
Dennis Singleton

Dennis Singleton has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.