Executive Summary
Rented GPU capacity is getting more expensive while the price of a token keeps falling. Nebius will raise on-demand rates for the H100, H200, B200 and B300 from 1 October, its second increase in three months, with the B300 up 21 percent to $9.50 a GPU hour. The memory tier moved further still, and Genoa memory rose about 41 percent.
The finding is that the floor under inference pricing is set by memory and power, not by the age of the silicon. Gartner still expects roughly a 90 percent fall in cost per token by 2030, and both things are true at once, because the tokens a GPU hour can produce climb faster than the rent does. Watch the spread between the two numbers. Either one read on its own will mislead you.
The cheapest way to buy AI compute is not getting cheaper. From 1 October, Nebius raises the on-demand rate for four Nvidia accelerators, and it is the second rise in three months. The number most teams track, cost per token, is still heading down.
Both of those live in the same market, which is why the gap between them is the story. If you rent GPU capacity this quarter, a higher GPU hour moves the break-even point for building a feature on your own hardware instead of paying per token to somebody else.
GPU rental prices have moved against the old instinct that hardware gets cheaper with age. The H100 launched in 2022, and it is more expensive to rent today than it was in May.
The new rate card is not a rounding error
Nebius published the change on its own pricing page. The H100 moves from $3.85 to $4.50 an hour, up 17 percent. The H200 rises 20 percent to $5.40. The B200 rises 19 percent to $8.50. The B300 takes the steepest step, up 21 percent to $9.50.
Add the three months since the previous change, and the B300 rate is up roughly 56 percent. The increases were not limited to accelerators. Nebius lifted the AMD EPYC Genoa vCPU rate 25 percent and the Genoa memory rate about 41 percent.
Trackers saw the pressure earlier. SemiAnalysis recorded H100 one-year contract prices climbing from about $1.70 a GPU hour in October 2025 to $2.35 by March 2026, close to 40 percent.
Memory sets the floor, not the silicon
Here is the mechanism. The newer the chip, the more high-bandwidth memory it carries, and memory has been the tightest part of the build, as memory market trackers keep recording. The Genoa memory rate rose about 41 percent in the same change, which is the memory squeeze showing up on the bill.
That is also why an aging H100 went up rather than down. The constraint is not how fast the chip is. It is how much memory ships with it, and how much power the rack can push through it.

Two curves, and both are real
The falling number is real too. Gartner expects cost per token to fall roughly 90 percent by 2030. That forecast is about throughput per dollar, not about rent. Serving stacks such as vLLM keep raising what a single GPU hour produces, through paged attention, speculative decoding and prefix caching.
So the rent rises and the token price falls, because the machine does more work in the hour it is rented. Our earlier look at the inference price war and cache economics covers the throughput side.
The specialist clouds are pricing to that scarcity. CoreWeave has been lifting contracted power, from 3.1 gigawatts at the end of 2025 to 4.2 gigawatts by August, and JPMorgan moved its rating to Overweight on stronger compute pricing. Nebius reported second-quarter revenue of $582.3 million, with AI cloud revenue up 514 percent to $574.9 million.
Three questions are worth asking before you sign a compute contract this quarter. Is your committed rate locked, or does it reset with the list price? What is your duty cycle, since an idle GPU hour bills the same as a busy one? And at what utilisation does owning the rack beat renting it?
Related reading. We traced how agent sandboxes fail when the runtime assumes the workload is polite.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.

[…] price of that compute is moving too. GPU rental rates are climbing while the token price is not. That analysis is here. And Anthropic’s filing shows what a decade of contracted compute looks like on a balance […]
[…] Related reading. For the power side of the same rack, see our piece on 800 volt DC delivery at one megawatt. For how memory cost shows up in rental rates, see GPU rental prices. […]
[…] reading. We traced the same pressure across the rising GPU hour and the falling token price, and across what a rack now has to earn per […]
[…] the wider cost picture, see why GPU rental prices are rising and why one lab bought inference instead of building […]
[…] For how that capacity is actually reaching production, see Vera Rubin NVL72 in production and our read on rising GPU rental prices. […]