Kimi K3 just gave Nvidia (NVDA) its second "DeepSeek moment." The first one made chip demand go up, not down.

Kimi K3 just gave Nvidia (NVDA) its second "DeepSeek moment." The first one made chip demand go up, not down.

Key points

  • Moonshot AI's Kimi K3, a 2.8 trillion-parameter open model out of Beijing, sent Nvidia (NVDA) down about 2.5% and Oracle (ORCL) down 6.3% on July 16.
  • The same "cheap AI kills chip demand" story ran in January 2025 with DeepSeek. Nvidia lost more than $500 billion in a day, then hyperscaler spending rose anyway.
  • Hyperscalers are on pace to spend around $650 billion on AI infrastructure in 2026, up 67% from last year.
  • Cheaper models tend to multiply usage faster than they cut costs, and Kimi K3's reasoning style burns more compute per answer, not less.

Moonshot AI, a Beijing startup founded in 2023, released a model called Kimi K3 this week. It reportedly matches or beats GPT-5.6 and Claude Fable 5 on several major benchmarks. Nvidia (NVDA) and Oracle (ORCL) stock both fell the same day, reviving a question Wall Street has asked before: if a Chinese lab can build a frontier model this cheaply, why does anyone need hundreds of billions of dollars in US chips?

Fair question. Wrong answer, going by what happened the last time it got asked.

What Kimi K3 actually is

Kimi K3 is a 2.8 trillion-parameter model built on a mixture-of-experts design, meaning it only activates a small slice of its full parameter count (16 of 896 "experts") for any single response, which is how it keeps inference costs down despite its size. It has a 1 million-token context window and is priced at roughly $3 per million input tokens and $15 per million output tokens through Moonshot's API, less than half what OpenAI charges for its flagship GPT-5.6 Sol ($5/$30 per million tokens) and less than a third of Anthropic's Claude Fable 5 ($10/$50 per million tokens). On Arena.ai's WebDev leaderboard, it opened at number one on July 16, ahead of both Claude Fable 5 and GPT-5.6.

Moonshot has not published a training cost for K3, and the open-weight release, which would let outside researchers actually check how it was built, is not due until July 27, so the efficiency claim can't be independently checked yet. The last time a Moonshot model made headlines for its price tag, a reported $4.6 million training cost for Kimi K2 Thinking circulated widely in November 2025. Moonshot's own chief executive, Yang Zhilin, said the number "is not official" and that training cost is "hard to quantify" when so much of it is research and experimentation rather than a single compute bill.

The DeepSeek precedent

DeepSeek set off the same argument in January 2025. It released a model called R1, claiming it cost around $6 million to train, a figure DeepSeek never fully substantiated to outside observers. Nvidia's shares fell enough that day to erase more than $500 billion in market value, a drop widely reported at the time as the largest single-day loss for any public company. The logic was identical to this week's: if frontier performance can be had for a few million dollars, the hundreds of billions being spent on American chips and data centers were about to look like a very expensive mistake.

They didn't. In the year that followed, Microsoft's and Google's cloud revenue grew 26% and 48%, and Google doubled its own capital spending just to keep up with backlogged demand. Major hyperscalers held or raised their AI capital spending guidance in the weeks after the DeepSeek shock, not cut it. Meta chief executive Mark Zuckerberg raised his own company's 2025 AI spending target within days of the news. Nvidia overtook Apple (AAPL) as TSMC's largest customer in 2025, a spot Apple had held for more than a decade.

Kimi K3 is triggering the same reflex, in the same week the chip sector is already nursing a rough month. It landed on July 16, the same day SK Hynix and Samsung-linked names were already sliding on a separate, Korea-specific selloff we've covered in detail here. Nvidia closed down about 2.5%, Oracle fell 6.3%, and AMD dropped another 5.3%, on top of a semiconductor sector that was already off double digits from its June peak. Some of that is Korea. Some of it, per multiple outlets, is Kimi K3 specifically reviving the "cheap AI kills chip demand" argument for the first time since DeepSeek.

Why cheaper AI raises compute demand instead of lowering it

There are two separate reasons the "less compute needed" read gets this backwards, and they compound each other.

The first is straightforward economics, sometimes called Jevons paradox: when something gets cheaper to use, people don't use the same amount for less money, they use a lot more of it. Microsoft chief executive Satya Nadella made this argument directly last year: "As AI gets more efficient and accessible, we will see its use skyrocket, turning it into a commodity we just can't get enough of." A cheaper Kimi K3 API doesn't just let existing users pay less. It puts frontier-level AI in reach of far more companies and use cases than could justify the old pricing, and in practice that has always meant total spend on compute goes up, not down.

The second reason is more specific to K3 itself. It's built as a reasoning model, meaning it works through intermediate steps before answering rather than producing a response in one pass. That style of model uses meaningfully more compute at the moment someone actually uses it, what the industry calls inference or test-time compute, than older, single-pass chatbots did, even if the initial training run was efficient. A cheap-to-train reasoning model doesn't reduce the total compute bill for a given task. It shifts where the bill shows up, from a one-time training run to millions of individually more expensive queries running every day.

Where the money is actually going

The spending data backs this up. Hyperscalers are on pace to spend roughly $650 billion on AI infrastructure in 2026, up 67% from last year, and combined capital spending across the largest cloud providers is projected near $1.15 trillion for 2025 through 2027, more than double what the same group spent in the prior three years. Microsoft and Oracle are now running capital spending at 45% and 57% of revenue, ratios closer to a utility company's than a software business's. We wrote about Meta and Amazon's own billion-dollar bets on this same demand thesis earlier this month, and it's the same story: nobody who actually writes the checks is treating cheaper models as a reason to slow down.

That is also why Meta's own "spare compute" announcement earlier this month got misread the same way K3 is being misread now. Meta didn't build Meta Compute because it has too much capacity sitting around forever, it built enough capacity that it can resell the temporary excess while still racing to add more. For the fuller picture of what all that spending is actually buying, from the chips themselves down to the power grid, we broke down the whole AI hardware stack here.

What could actually go wrong

None of this means chip stocks are a one-way trade. The honest risks are real and worth naming. The sector rallied roughly 130% over the twelve months before this selloff started, which left almost no room for anything less than a flawless run of earnings, and several genuinely strong reports this earnings season got sold anyway. Oracle in particular is funding a large share of its AI buildout with debt against a backlog that hasn't turned into cash flow yet, which is a real financing risk if AI revenue growth disappoints, separate from anything Kimi K3 does. And the memory supply chain that feeds every AI chip still runs through two Korean companies working through their own, unrelated, leverage problems.

Expensive and richly valued is not the same claim as unnecessary. Kimi K3, like DeepSeek before it, backs up the first one. It does nothing for the second, and the last eighteen months are a fairly clean natural experiment showing why.

Sources

Stock prices are as of market close, July 16, 2026, unless noted otherwise. This is general market commentary and not investment advice. Always do your own research and consider speaking with a licensed financial professional before making any investment decision.

Frequently asked questions

What is Kimi K3 and why did it move chip stocks?

Kimi K3 is a 2.8 trillion-parameter open-weight AI model from Chinese startup Moonshot AI, released July 16, 2026, that reportedly matches or beats GPT-5.6 and Claude Fable 5 on several benchmarks at a fraction of the price. Its release revived fears that cheaper Chinese AI models reduce the need for expensive US chips, and Nvidia (NVDA) closed down about 2.5% and Oracle (ORCL) fell 4.3% that same day.

Is this the same as the DeepSeek moment from 2025?

It follows a near-identical pattern. DeepSeek's R1 model, released in January 2025 with a claimed $6 million training cost, wiped out more than $500 billion of Nvidia's market value in a single day, at the time the largest one-day loss for any company in stock market history. Hyperscaler AI spending rose in the year that followed rather than falling.

Does a cheaper AI model actually reduce demand for chips like Nvidia's?

Historically, no. Cheaper AI tends to expand total usage faster than it lowers the per-use cost, an effect economists call Jevons paradox, and reasoning models like Kimi K3 use more compute at the moment they're actually used, even when the training run itself was efficient. Hyperscaler AI infrastructure spending is on pace for around $650 billion in 2026, up 67% from 2025.

Has Moonshot released Kimi K3's training cost or open weights?

Not yet as of this writing. Moonshot has said the full open-weight release is expected by July 27, 2026, and has not published an official training cost for K3. Its prior model, Kimi K2 Thinking, had a widely reported $4.6 million training figure that Moonshot's own CEO said was not an official number.

More on AAPL and AMD

Dennis Singleton
Dennis Singleton

Dennis Singleton has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.