Photo by Yuriy Vertikov on Unsplash. Source: https://unsplash.com/photos/yZpcIxmkxVE (Unsplash License).

Executive Summary

Memory has become the biggest moving cost inside an AI rack. On Morgan Stanley’s teardown of the two most recent generations, memory goes from roughly 9 percent of a rack’s material cost to about 26 percent, and the memory line alone rises from about $374,000 to just over $2 million. The GPU’s share of the same bill falls toward half.

Nvidia’s answer has been arithmetic rather than engineering. It halved the CPU side memory on its flagship rack and halved the memory and the storage on a desktop system to hold a price point, while the larger version of that system rose to $6,950. The verdict for anyone planning capacity is blunt. A capex model built on 2025 rack pricing is already wrong, and a 17 percent rack increase adds more than $5 billion to a single gigawatt of AI capacity.

The constraint on AI capacity has moved. For three years the scarce part of a cluster was the accelerator, and the queue was for GPUs. In 2026 the part that sets the price is AI server memory, the DRAM and high bandwidth memory wrapped around each chip. The cause is wafer arithmetic rather than a sudden jump in demand.

High bandwidth memory is built by stacking DRAM dies and joining them with through silicon vias. It carries more bandwidth than any conventional module, and it consumes roughly four times the wafer area per usable gigabyte. When a single accelerator needs hundreds of gigabytes of it and a rack carries dozens of accelerators, the wafer area adds up fast. Analysts put HBM at roughly a fifth of global DRAM wafer capacity this year, up from under a tenth in 2024. Every wafer moved to HBM is a wafer not making the server DIMMs that everything else still needs.

The Bill of Materials Flipped in One Generation

The contract prices show the squeeze. Conventional DRAM contract prices rose 90 to 95 percent quarter on quarter in the first quarter of 2026, then a further 58 to 63 percent in the second, with Counterpoint Research tracking the same pace across DRAM, NAND and HBM. TrendForce now forecasts another 13 to 18 percent in the third quarter, a slowdown driven by buyers rationing rather than by supply improving.

On the rack, that lands as a structural shift. Morgan Stanley’s estimate puts the memory content of a Vera Rubin NVL72 rack above $2 million inside a total near $7.8 million, up more than 400 percent from the same line on the prior generation. Memory is no longer a rounding item next to the silicon, and GPU silicon is down to roughly half of the bill from more than 80 percent.

The Cut Is the Tell, Not the Increase

The most revealing move is not a price rise. Nvidia reportedly halved the default CPU side memory on a Vera Rubin rack, taking the LPDDR5X modules from 192GB to 96GB each, which cuts rack level LPDDR5X from roughly 54 terabytes to 28. GF Securities puts the saving near $586,000 per rack against $1.2 million, and TrendForce’s read is that preliminary supplier allocation covered only about 60 percent of Nvidia’s low power DRAM demand. The parts to fill every slot do not exist, so the shipping configuration got leaner.

The desktop side tells the same story in a smaller frame. Nvidia introduced a 64GB version of its GB10 workstation at $4,999, and the 128GB version rose to $6,950, up from a $3,999 launch price. The detail worth keeping is that memory bandwidth is unchanged at 273 gigabytes per second, which means Nvidia used lower capacity modules rather than fewer of them. This is a supply workaround, not a slower machine.

What to Do With a Number That Keeps Moving

Three practical moves follow. Lock supply rather than price, because long term agreements with memory suppliers, or with the manufacturers that hold allocation, are the only shield TrendForce identifies. Re-baseline capex now, since a rack price rise of 15 to 17 percent is a board level variance across a large deployment. And right-size memory configurations, using Nvidia’s own cut as the template, shipping what the workload needs today and leaving room to add capacity later.

The timing matters. Deloitte expects meaningful new memory fab capacity only in 2029 or 2030, and The Register reports Micron telling the market that RAM supply gets worse before it improves. Gartner’s nearer view has tightness lasting through at least the first half of 2027. Plan on expensive memory for the life of the next cluster, not for one more quarter.

Diagram showing memory rising from 9 to 26 percent of rack material cost while GPU silicon falls from over 80 percent to about 50 percent, with key DRAM price and DGX Spark figures.
Memory more than doubled its share of rack material cost in one generation. Sources, Morgan Stanley estimate via Tom’s Hardware, TrendForce contract surveys, Deloitte.

Related reading. For the power side of the same rack, see our piece on 800 volt DC delivery at one megawatt. For how memory cost shows up in rental rates, see GPU rental prices.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

One thought on “Memory Now Sets the Price of an AI Rack, Not the GPU”

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.