Photo by Brecht Corbeel on Unsplash. Source: https://unsplash.com/photos/bright-light-illuminates-a-microchip-on-a-dark-circuit-board-cQ0OdrlUPtw (Unsplash License).

New accelerator generation, and the serving software was already ready. vLLM runs on Nvidia’s Vera Rubin NVL72 while the rack is still in early access.

  • 7.8x the per-GPU throughput of GB200 NVL72 on the AgentX agentic serving benchmark, and up to 3.7x vision-language throughput over GB300 NVL72 in MLPerf.
  • Per GPU, NVFP4 inference rises from 10 to 50 PFLOPS and HBM bandwidth rises from 7.94 to 19.16 terabytes per second.
  • The gain comes from memory, not raw math. Decode is memory-bound, and the team split the mixture-of-experts weights across Rubin’s two locality domains so each processor reads local memory.
  • Rubin kept the Blackwell instruction set, so kernels ported on day zero. Support shipped for DeepSeek, Kimi, GLM and MiniMax before general availability.
  • Blunt caveat. The 7.8x is one benchmark on one model family, and HBM4 supply is tight.

Read the full analysis.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.