The unit that decides an inference bill is the rack, not the chip. CoreWeave put NVIDIA Vera Rubin NVL72 into production on September 30 and published the first customer-run numbers for it.
- Vera Rubin NVL72 carries 72 Rubin GPUs and 36 Vera CPUs in a single rack.
- Cognition, the lab behind the Devin coding agent, measured up to 4.8 times the total token throughput of its own GB200 NVL72 baseline on coding work.
- On reinforcement learning it saw 3.8 times more output tokens.
- CoreWeave reported 10 times the token throughput per megawatt on DeepSeek R1 at matched interactivity.
- One Vera CPU rack holds 11,264 cores, enough for more than 11,000 concurrent sandbox environments.
- The caveat is that these are vendor-run figures on one lab’s workloads, with no independent benchmark published yet.
The takeaway for operators is that throughput per megawatt now ties a rack to a power contract, and the power contract is the long pole in every build.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.
