Executive Summary
NVIDIA’s flagship data center products are no longer servers or accelerator cards. They are racks. In GB200 and GB300 NVL72 the GPUs are soldered to compute trays, joined by NVLink into a single domain, and sold as one unit. There is no smaller granularity, because the interconnect does not work at smaller scale. The practical minimum order is one rack, which is 72 GPUs.
That changes what a buyer is actually choosing. It is no longer a GPU count at a unit price. It is a power envelope, a cooling plant, a depreciation schedule running against an annual product cadence, and a financing structure. Morgan Stanley puts a GB300 NVL72 rack near 4 million dollars and a Vera Rubin NVL72 rack near 7.8 million, and NVIDIA publishes neither figure itself. A single rack draws 120 to 230 kilowatts against 20 to 30 kilowatts for the air-cooled racks most facilities were built around.
The integration boundary moved twice in three years, and it keeps moving up a level. In the older HGX form factor, eight GPUs sat on one board inside a normal server. In the NVL72 rack, the GPU is soldered onto a tray, eighteen trays fill the rack, and the reference architecture describes all 72 GPUs functioning as one multi-GPU unit of compute.

The unit of purchase moved from the GPU to the rack
The GB200 NVL72 announcement introduced 72 Blackwell GPUs and 36 Grace CPUs joined by fifth-generation NVLink at 130 terabytes per second of aggregate bandwidth. The GB300 NVL72 followed with 72 Blackwell Ultra GPUs. Vera Rubin moved the unit again, to a POD of five purpose-built racks at 1,152 GPUs.
Because the domain does not subdivide, a buyer cannot start with four GPUs and scale. They buy a rack or a POD, from an OEM, pre-integrated. Dell, HPE, Lenovo, Supermicro, ASUS, Foxconn, GIGABYTE and others take them to market, and HPE lists GB200 NVL72 by HPE as a rack-scale system with a single rack power figure rather than a GPU part number.
NVIDIA does not publish list prices. Every figure below is analyst or press reporting, and is labeled as such.
Depreciation now runs against a twelve month product cadence
Neoclouds depreciate physically identical hardware over different lives. CoreWeave uses six years, Lambda five, Nebius four. Meanwhile NVIDIA ships a new rack generation roughly annually, Blackwell in 2024, Blackwell Ultra in 2025, Vera Rubin in 2026, Rubin Ultra in 2027.
A six-year schedule therefore assumes the asset stays economic across three or four replacement generations. CoreWeave reported 1,510 million dollars of adjusted EBITDA at a 59 percent margin alongside 1,393 million dollars of depreciation and amortisation in its second quarter, so roughly 92 percent of that EBITDA is depreciation added back on the GPU fleet. Financing has moved into the product as well. NVIDIA’s August 2026 tie-ups with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR aim to mobilise over 500 billion dollars of third-party capital, though NVIDIA describes them as memorandums subject to final agreements.
The counter-argument is real. CoreWeave holds A100 service contracts running to 2029 on a GPU launched in 2020, which would be a nine-year economic life if it holds.
You cannot upgrade one GPU, so the rack is the renewal unit
Because the 72 GPUs form one NVLink domain and are soldered to trays, there is no swapping one accelerator for a faster one. The options are replacing the whole rack or running a new generation beside the old one. That is why Supermicro now expresses its blueprint as multiplying a 1,152-GPU unit sold against a fixed power envelope from five megawatts to a gigawatt.
The lock-in is architectural rather than commercial. The domain is proprietary, the software stack is CUDA, and the rack cannot be partially filled or mixed with another vendor’s accelerators. AMD is attacking exactly that with Helios, a 72-GPU rack on open UALink fabric, and claims more HBM capacity and scale-out bandwidth than a Vera Rubin NVL72.
Downstream the arithmetic is unchanged but the exposure is larger. Neoclouds sell GPU-hours rather than racks, and Lambda lists HGX B200 from 9.86 dollars per GPU-hour on a small cluster. CoreWeave carried 35.1 billion dollars of total debt at 30 June 2026 with 640 million dollars of quarterly interest expense against 128 million dollars of adjusted operating income.
Three questions for the next refresh cycle. Does your depreciation life survive three more generations, or is it borrowed from a slower market? Is your power envelope provisioned for 142 kilowatts a rack or for what the next rack will need? And when the upgrade arrives, is your plan to replace racks or to run generations side by side in the same hall?
For how that capacity is actually reaching production, see Vera Rubin NVL72 in production and our read on rising GPU rental prices.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.
