Photo by ThisisEngineering on Unsplash. Source: https://unsplash.com/photos/person-using-computer-on-table-_5N2rhCqtC0 (Unsplash License).

Executive Summary

Rented GPU capacity is getting more expensive while the price of a token keeps falling. Nebius will raise on-demand rates for the H100, H200, B200 and B300 from 1 October, its second increase in three months, with the B300 up 21 percent to $9.50 a GPU hour. The memory tier moved further still, and Genoa memory rose about 41 percent.

The finding is that the floor under inference pricing is set by memory and power, not by the age of the silicon. Gartner still expects roughly a 90 percent fall in cost per token by 2030, and both things are true at once, because the tokens a GPU hour can produce climb faster than the rent does. Watch the spread between the two numbers. Either one read on its own will mislead you.

The cheapest way to buy AI compute is not getting cheaper. From 1 October, Nebius raises the on-demand rate for four Nvidia accelerators, and it is the second rise in three months. The number most teams track, cost per token, is still heading down.

Both of those live in the same market, which is why the gap between them is the story. If you rent GPU capacity this quarter, a higher GPU hour moves the break-even point for building a feature on your own hardware instead of paying per token to somebody else.

GPU rental prices have moved against the old instinct that hardware gets cheaper with age. The H100 launched in 2022, and it is more expensive to rent today than it was in May.

The new rate card is not a rounding error

Nebius published the change on its own pricing page. The H100 moves from $3.85 to $4.50 an hour, up 17 percent. The H200 rises 20 percent to $5.40. The B200 rises 19 percent to $8.50. The B300 takes the steepest step, up 21 percent to $9.50.

Add the three months since the previous change, and the B300 rate is up roughly 56 percent. The increases were not limited to accelerators. Nebius lifted the AMD EPYC Genoa vCPU rate 25 percent and the Genoa memory rate about 41 percent.

Trackers saw the pressure earlier. SemiAnalysis recorded H100 one-year contract prices climbing from about $1.70 a GPU hour in October 2025 to $2.35 by March 2026, close to 40 percent.

Memory sets the floor, not the silicon

Here is the mechanism. The newer the chip, the more high-bandwidth memory it carries, and memory has been the tightest part of the build, as memory market trackers keep recording. The Genoa memory rate rose about 41 percent in the same change, which is the memory squeeze showing up on the bill.

That is also why an aging H100 went up rather than down. The constraint is not how fast the chip is. It is how much memory ships with it, and how much power the rack can push through it.

Chart showing per-token inference prices falling while the rented GPU hour rises, with the Nebius on-demand rate card up 17 to 21 percent on 1 October 2026.
Per-token prices keep falling while the rented GPU hour climbs. Nebius raises four accelerator tiers on 1 October.

Two curves, and both are real

The falling number is real too. Gartner expects cost per token to fall roughly 90 percent by 2030. That forecast is about throughput per dollar, not about rent. Serving stacks such as vLLM keep raising what a single GPU hour produces, through paged attention, speculative decoding and prefix caching.

So the rent rises and the token price falls, because the machine does more work in the hour it is rented. Our earlier look at the inference price war and cache economics covers the throughput side.

The specialist clouds are pricing to that scarcity. CoreWeave has been lifting contracted power, from 3.1 gigawatts at the end of 2025 to 4.2 gigawatts by August, and JPMorgan moved its rating to Overweight on stronger compute pricing. Nebius reported second-quarter revenue of $582.3 million, with AI cloud revenue up 514 percent to $574.9 million.

Three questions are worth asking before you sign a compute contract this quarter. Is your committed rate locked, or does it reset with the list price? What is your duty cycle, since an idle GPU hour bills the same as a busy one? And at what utilisation does owning the rack beat renting it?

Related reading. We traced how agent sandboxes fail when the runtime assumes the workload is polite.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

5 thoughts on “GPU Rental Prices Are Going Up. The Token Price Is Not.”

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.