A person installing a graphics card on a computer. Photo by Samsung Memory on Unsplash. Source: https://unsplash.com/photos/a-person-holding-a-device-iRY8lnq0xb0 (Unsplash License).

The unit that decides an inference bill is the rack, not the chip. CoreWeave put NVIDIA Vera Rubin NVL72 into production on September 30 and published the first customer-run numbers for it.

  • Vera Rubin NVL72 carries 72 Rubin GPUs and 36 Vera CPUs in a single rack.
  • Cognition, the lab behind the Devin coding agent, measured up to 4.8 times the total token throughput of its own GB200 NVL72 baseline on coding work.
  • On reinforcement learning it saw 3.8 times more output tokens.
  • CoreWeave reported 10 times the token throughput per megawatt on DeepSeek R1 at matched interactivity.
  • One Vera CPU rack holds 11,264 cores, enough for more than 11,000 concurrent sandbox environments.
  • The caveat is that these are vendor-run figures on one lab’s workloads, with no independent benchmark published yet.

The takeaway for operators is that throughput per megawatt now ties a rack to a power contract, and the power contract is the long pole in every build.

Read the full analysis

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.