New accelerator generation, and the serving software was already ready. vLLM runs on Nvidia’s Vera Rubin NVL72 while the rack is still in early access.
- 7.8x the per-GPU throughput of GB200 NVL72 on the AgentX agentic serving benchmark, and up to 3.7x vision-language throughput over GB300 NVL72 in MLPerf.
- Per GPU, NVFP4 inference rises from 10 to 50 PFLOPS and HBM bandwidth rises from 7.94 to 19.16 terabytes per second.
- The gain comes from memory, not raw math. Decode is memory-bound, and the team split the mixture-of-experts weights across Rubin’s two locality domains so each processor reads local memory.
- Rubin kept the Blackwell instruction set, so kernels ported on day zero. Support shipped for DeepSeek, Kimi, GLM and MiniMax before general availability.
- Blunt caveat. The 7.8x is one benchmark on one model family, and HBM4 supply is tight.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.
