Executive Summary
Thinking Machines Lab signed a $65 million annual agreement with Crusoe on September 23 to run production inference for its open models. The work runs on a dedicated deployment of NVIDIA HGX B200 systems connected with NVIDIA Quantum-2 InfiniBand networking. Crusoe operates the cluster. The lab pays for the output.
The size of the deal is not the signal. The milestone behind it is. Crusoe Managed Inference crossed $100 million in contracted annual recurring revenue less than a year after launch, and all of it is signed. Training capacity and inference capacity have become separate purchases. Only one of them still requires the buyer to own the stack.
A managed inference endpoint is an odd product when you write it down. You do not get GPUs. You get a service that returns tokens, backed by a cluster the vendor owns, tunes, and patches.
Crusoe calls its version a Tailored Deployment. The customer gets a dedicated, benchmarked, SLA-backed endpoint. Crusoe runs the cluster directly. The customer sends traffic and pays for what it serves. There is no rack to cable and no serving stack to upgrade.
Thinking Machines Lab signed up for that on September 23. The lab will serve production traffic for its Inkling models, GLM 5.2 and 5.3, and its own fine-tuned variants on a dedicated deployment of NVIDIA HGX B200 systems connected with NVIDIA Quantum-2 InfiniBand networking. The agreement runs $65 million a year.
Myle Ott, the lab’s ML infrastructure lead, said the managed service met the standards the lab holds its own systems to. The harder reading is simpler. A lab that prizes infrastructure quality rented the serving layer.

The $65 million is the wrong number to watch
Breadth matters as much as any single logo. Crusoe announced a multi-year agreement with Perplexity the week before this deal, then closed a $3.9 billion Series F round. Managed serving is the fastest moving line in a business that started out renting GPU capacity.
The number Crusoe attached to the deal matters more than the deal. Crusoe Managed Inference passed $100 million in contracted annual recurring revenue less than a year from launch.
Contracted is the word doing the work. Not trials, not evaluations, not cloud credits. Signed annual commitments. That is a demand signal for managed serving as a category, not for one lab’s budget.
$65 million is a rounding error in a market where single capacity leases run into gigawatts. As evidence that labs will buy inference instead of building it, it is a different thing entirely.
Inference is becoming its own purchase order
Training capacity arrives through direct chip deals, power contracts, and multi-year commitments to a physical site. Inference capacity needs none of that. The workload is elastic, the pricing is per token, and the switching cost is low.
The buying pattern follows. Labs lock training down with capital. They shop inference among the neoclouds on price, latency, and reliability. Crusoe competes on that list against CoreWeave, Lambda, Together AI, and the hyperscalers’ own serving products. The open serving engines underneath are similar, and vLLM is not the differentiator. Running it well is.
That is the point of a $65 million contract. Running your own inference stack stopped being the default. It became a choice, and an expensive one.
Related reading. Anthropic leased 2.16 gigawatts in Australia for inference alone, and AWS made GPU aware inference routing an EKS add on.
Three questions for your own stack. Do you know what share of your AI bill is training and what share is inference. If your serving team disappeared tomorrow, could you describe the endpoint you would rent. And is the capacity you depend on contracted, or is it best effort.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.

[…] supplies orchestration, operations, and the developer contract. It is the same split we saw in the managed inference deals from earlier this year, where a model lab hands operations to a specialist and keeps the API in […]