Executive Summary
CoreWeave put NVIDIA Vera Rubin NVL72 into production and published the first customer-run numbers for it. Cognition, the lab behind the Devin coding agent, measured up to 4.8 times the total token throughput of its own GB200 NVL72 baseline on coding work, and 3.8 times more output tokens on reinforcement learning. CoreWeave added a figure of its own, 10 times the token throughput per megawatt on the DeepSeek R1 reasoning model at matched interactivity.
The finding is that the number that decides an inference purchase has moved off the chip. Peak throughput stopped predicting cost when serving became memory-bound, so what a buyer can compare now is tokens served per megawatt, and that moves the rack count and the power contract together. The caveat is real. These are vendor-run tests, on one lab’s workloads, comparing new silicon with the generation it replaces.

Rented AI capacity is sold by the GPU hour. The number that decides whether the deal works is how many tokens that hour produces. CoreWeave put NVIDIA’s newest rack into production on September 30 and gave that second number a public baseline.
CoreWeave runs the platform. Vera Rubin NVL72 is one rack that carries 72 Rubin GPUs and 36 Vera CPUs, joined by NVLink and fronted by BlueField-4 DPUs. Inference has two halves, the prefill that reads the prompt and the decode that writes the answer, and keeping both inside one high-bandwidth rack makes the cache transfer cheap.
{FIGA}
The measured gain is throughput per megawatt, not per chip
The comparison Cognition published is against the GB200 NVL72 rack it already ran. On its SWE-2 coding workload it saw up to 4.8 times the total token throughput, and on reinforcement learning it measured 3.8 times the output token throughput. Both matter because agentic coding holds long contexts at high concurrency, so every step waits on the one before it.
CoreWeave’s own headline is the number to read twice. It claims 10 times the token throughput per megawatt over GB200 NVL72 on DeepSeek R1 at matched interactivity. That qualifier is doing the work. Throughput only compares when latency is held equal, and a per-megawatt figure folds power into the cost of a single token.
The floor under the whole exercise is memory and power, not raw compute. We made that case when token prices halved in a single night and the bill moved into the cache. NVIDIA Dynamo, the open source serving framework behind CoreWeave’s managed inference, is the software half of the same push.
The agent sandbox is why the CPU half matters
Not all of the gain comes from the GPU. CoreWeave will serve NVIDIA Vera, the CPU NVIDIA built for agents, as bare metal at rack scale. One Vera rack carries 128 processors and 11,264 cores, enough for more than 11,000 concurrent isolated environments.
That number is the point. An agentic workload does not run one model and stop. It spawns thousands of short-lived sandboxes to execute code, call tools and run tests. CoreWeave measured more than 3 times faster sandbox start times on Vera, and a 1.7 times gain on Terminal-Bench. The bottleneck it attacks is the rate at which isolated environments can be created.
We tracked the same boundary from the other side this week. DeepSeek published a catalog of agents that escaped their sandbox, and Docker moved an agent’s permissions into the image it runs. The CPU that hosts those sandboxes is now part of the security story, not just the serving story.
What an operator should take from a vendor benchmark
The numbers are directional, not a price sheet. They come from a vendor, run on one partner’s workloads, and compare new silicon against the part it replaces. Nobody has published an independent, mixed-workload, per-megawatt comparison yet, and that is the number procurement will eventually ask for.
The durable claim is narrower and more useful. Throughput per megawatt is the metric that ties a rack to a power contract, and the power contract is the long pole in every AI build right now. That reframes the purchase from a chip decision into a facility decision.
Three questions for your own plan. Which of your workloads are throughput-bound and which are memory-bound. What is your current tokens-per-megawatt on the hardware you already run. And how many concurrent sandboxes your agent platform genuinely needs at peak.
Related reading. We looked at why the cold start was the model, not the container, and at why AI factories now buy the battery before the chip.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.
