Executive Summary
vLLM runs on Nvidia’s Vera Rubin NVL72 while the rack is still in early access, and the headline is a 7.8x gain in throughput per GPU over GB200 NVL72 on the AgentX agentic serving benchmark. A second result puts vision-language throughput at up to 3.7x over GB300 NVL72 in MLPerf. These are nightly container builds, not a finished release.
The part that matters is where the gain comes from. NVFP4 inference per GPU rose five times, and raw math did not do the work. Memory bandwidth rose to 19.16 terabytes per second from 7.94, and the vLLM team rebuilt the mixture-of-experts path to read weights from the nearest memory domain. For a buyer, the case for the generation after Blackwell rests on memory bandwidth per dollar and per watt, not on FLOPS alone.
New accelerators usually arrive with a months-long software gap. The kernels do not exist, the serving stack does not support the model family, and the first honest numbers land quarters later. vLLM closed that gap on Vera Rubin before most operators have seen the rack, and the reason traces back to a decision NVIDIA made in the silicon.
The Rubin family kept the Blackwell instruction set, so the kernels carried over
NVIDIA built Vera Rubin on the Blackwell architecture family. That one choice means vLLM’s existing Blackwell kernels run on the new hardware. The vLLM team published support on 9 October for DeepSeek, Kimi, GLM and MiniMax without waiting for general availability.
The tooling filled in the rest. FlashInfer 0.7.0 supplies Rubin-tuned attention, GEMM and MoE kernels, and the team tuned a sparse attention prefill kernel for the new part. Engineers from NVIDIA, Inferact and Red Hat ran the port, and the project now publishes nightly containers on Docker Hub.
This is the inverse of the normal pattern. Teams usually budget a software lag into every hardware refresh, and that lag is what makes a first-year purchase hard to justify. When the instruction set carries forward, the lag shrinks and the new rack starts earning sooner.
The throughput gain sits in memory, not in raw math
Start with the per-GPU deltas against GB200 NVL72. NVFP4 inference rises from 10 to 50 PFLOPS. HBM moves from HBM3e to HBM4, and bandwidth rises from 7.94 to 19.16 terabytes per second. Bidirectional NVLink bandwidth rises from 1.8 to 3.0 terabytes per second.
The FLOPS number is the one the marketing leads with, and it is not the one that produced the throughput. Decode is memory-bound. A model step spends its time moving weights and the key-value cache, not executing matrix multiplies. That is why a 2.4x memory gain can drive a larger end-to-end result than the arithmetic alone would suggest.
The vLLM team leaned into it. Rubin exposes two locality domains per GPU, a feature the CUDA programming guide describes as pairing a group of streaming multiprocessors with its own slice of HBM. The team split the mixture-of-experts weights by column and pinned each half to one domain, so each processor reads only local memory. On MiniMax M3 shapes that returned about 1.2x on the MoE layer by itself. The locality domain documentation covers the mechanics.
What this changes in a cluster budget
Take 7.8x per GPU at face value and the arithmetic gets uncomfortable for the status quo. The same agent traffic needs roughly an eighth of the accelerators, which cuts rack count, power draw and cooling. That is the cost-per-token figure every AI infrastructure team is now asked to defend.
Two caveats belong next to that number. The 7.8x comes from one benchmark, agentic serving at matched interactivity, on a single model family. And HBM4 lands in a market where memory supply is already tight, so the bandwidth that drives the gain is also the component most likely to constrain volume. The Vera Rubin NVL72 specification page and the GB200 NVL72 specification page carry the per-rack figures, and the NVIDIA developer breakdown explains the six-chip design.
Three questions are worth asking of your own fleet. Which of your serving workloads actually scale with memory bandwidth rather than arithmetic. How much of your current GPU bill is spent on memory stalls that never show up in a utilization graph. And whether your next refresh still plans for a software lag that may no longer exist.
Related reading. The cheapest inference on the market is a classifier, not a chat model, which changes where a serving stack spends its GPUs. The week’s roundup traces how every constraint in AI infrastructure now lands as a line item.

Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.
