Business

Nvidia Is Raising Prices 15%. OpenAI Just Cut Prices 33%. The AI Economy Is Splitting in Two.

CRAZE CRAZE Summary 3 things to know
  • Nvidia raises prices 15% due to memory shortage, preserving margins while server makers pass costs to cloud giants.
  • Memory now 25% of AI server costs; DRAM prices nearly doubled as Samsung, SK hynix, Micron lock supply until 2029-2030.
  • OpenAI cut GPT-5.6 Sol output token prices 33% as competition intensifies; hardware suppliers raise prices because component costs rise.
Jeff Editorial | · 5 min read
Nvidia Is Raising Prices 15%. OpenAI Just Cut Prices 33%. The AI Economy Is Splitting in Two.

On August 23, Bloomberg reported that Nvidia notified key clients of price hikes exceeding 15% for systems based on the Grace Blackwell and next-generation Vera Rubin architectures. The increases apply to shipments starting early next year. The trigger is memory supply shortages—HBM and server DRAM prices have surged, and Nvidia, despite its 75% gross margins, can no longer absorb the cost.

It's a striking contrast. Model providers are cutting prices. Hardware suppliers are raising them. The AI value chain is splitting into two economies.

Memory Now Eats 25% of Every AI Server

Memory has become the bottleneck that even Nvidia cannot outrun. DRAM contract prices jumped 93-98% in Q1 2026 alone. AI server DRAM is expected to quadruple in price over 2026. Deloitte estimates memory now accounts for roughly one-quarter of the bill of materials for high-end AI server racks—up from just 5% two years ago.

Samsung, SK hynix, and Micron control nearly 90% of global output, and their pricing power has become absolute. Even as they expand capacity, new fabs won't come online until 2029 or 2030. For the next three years, supply is locked.

The cost math is brutal. One estimate puts the Vera Rubin NVL72 rack at roughly $7.8 million, with memory accounting for about $2 million—a 435% jump from GB300 systems. A 1-gigawatt data center faces at least $50 million in additional construction costs from this round of price hikes alone.

Nvidia is not absorbing that difference. Server manufacturers have already passed the cost increases to Microsoft, Google, and Oracle. Nvidia's move ensures its own margins stay intact while the rest of the supply chain absorbs the shock.

OpenAI Is Cutting Prices. Nvidia Is Raising Them.

The contrast could not be sharper. On August 21, OpenAI announced it is cutting GPT-5.6 Sol API prices by over 20% for the next three months—output tokens drop from $30 to $20 per million. Google cut Gemini 3.7 Flash pricing to half of its predecessor. Meanwhile, Apple has raised Mac and iPad prices by nearly 20% due to memory costs, and Amazon lifted Echo Dot prices by 60%.

Model providers are cutting prices because their marginal cost per token is dropping. Hardware suppliers are raising prices because their marginal cost per component is rising. The two trends are moving in opposite directions, and they are both right—for where they sit.

Nvidia Is Raising Prices 15%. OpenAI Just Cut Prices 33%. The AI Economy Is Splitting in Two.
Nvidia is raising AI server prices 15% on memory costs. OpenAI just cut GPT-5.6 Sol by 33%. The AI economy is splitting in two.

$50 Million More Per Data Center—and Rising

Morgan Stanley has already coined a term for this: "chipflation." The memory price surge began as an AI infrastructure bottleneck—now it is spreading into device affordability, cloud costs, and even broader inflation.

The impact is measurable. Software and computer accessories, which usually trend cheaper over time, were up 14.5% year-over-year in May. Bloomberg Economics estimates the memory squeeze will add 0.4 percentage points to headline inflation before it eases. Data center AI capital expenditure is projected to reach $650 billion by 2026, up from $217 billion in 2024.

What It Means

The AI economy is not one market. It is three: chip suppliers, cloud providers, and model developers. Each faces different cost structures and competitive pressures. Nvidia is raising prices because it can—demand is inelastic, and memory is the bottleneck. OpenAI is cutting prices because it must—competition is intensifying, and developers have alternatives.

The divergence is not a contradiction. It is a signal that the AI value chain is maturing—and the profits are moving to whichever layer holds the scarce resource. Right now, that is memory chips. Next year, it might be something else.


P.S. The kicker: OpenAI doubled GPT-5.5 prices from $1.25 to $5 per million input tokens earlier this year, while Nvidia was cutting inference costs 35x. The pricing trends in AI are not linear—they are moves and countermoves in a market that has not yet found equilibrium. The only certainty is that the cost structure will keep shifting.


Frequently Asked Questions

Q: Why is Nvidia raising AI server prices?

A: Nvidia is raising prices by more than 15% on AI servers due to soaring memory costs. DRAM contract prices jumped 93-98% in Q1 2026 alone, and memory now accounts for roughly 25% of the bill of materials for high-end AI server racks—up from just 5% two years ago .

Q: How much is Nvidia raising prices?

A: Nvidia notified major customers of price hikes exceeding 15% for systems based on Grace Blackwell and next-generation Vera Rubin architectures . The increases apply to shipments starting early next year.

Q: Why is memory getting so expensive?

A: AI demand for HBM (High Bandwidth Memory) and server DRAM has surged, while supply remains constrained. Samsung, SK hynix, and Micron control nearly 90% of global output, and new fabs won't come online until 2029 or 2030 . Deloitte estimates AI server DRAM prices will quadruple over 2026 .

Q: How does this compare to OpenAI's price cuts?

A: OpenAI cut GPT-5.6 Sol API prices by 33% (output tokens from $30 to $20 per million) on August 21 . Google also cut Gemini 3.7 Flash pricing in half . The AI economy is splitting: hardware suppliers are raising prices due to cost pressure, while model providers are cutting prices due to competition .

Q: What is "chipflation"?

A: "Chipflation" is a term coined by Morgan Stanley to describe the inflationary impact of memory price surges . The memory squeeze, which began as an AI infrastructure bottleneck, is now spreading into device affordability, cloud costs, and even broader inflation . Bloomberg Economics estimates it will add 0.4 percentage points to headline inflation before easing .

Q: How much does memory cost in a typical AI server?

A: One estimate puts the Vera Rubin NVL72 rack at roughly $7.8 million, with memory accounting for about $2 million—a 435% jump from GB300 systems . A 1-gigawatt data center faces at least $50 million in additional construction costs from this round of price hikes alone .

Q: Who is paying the higher prices?

A: Server manufacturers have already passed the cost increases to Microsoft, Google, and Oracle . Ultimately, cloud customers and end users bear the cost through higher cloud service and device prices.

Q: Are other products affected by memory price hikes?

A: Yes. Apple has raised Mac and iPad prices by nearly 20% due to memory costs . Amazon lifted Echo Dot prices by 60% . Software and computer accessories, which usually trend cheaper over time, were up 14.5% year-over-year in May .

Q: How long will the memory shortage last?

A: SK hynix CEO has warned that demand could exceed supply until 2030 or beyond . Gartner predicts the supply shortage will continue at least through the first half of 2027 . New fabs currently under construction won't begin production until 2029 or 2030 .

Q: What does this mean for the AI industry?

A: The AI value chain is maturing—and profits are moving to whichever layer holds the scarce resource. Currently, that is memory chips. The divergence between hardware price hikes (Nvidia) and software price cuts (OpenAI) signals that different parts of the AI economy face different cost structures and competitive pressures .

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article