Executive Summary

Thinking Machines Lab signed a $65 million annual agreement with Crusoe on September 23 to run production inference for its open models. The work runs on a dedicated deployment of NVIDIA HGX B200 systems connected with NVIDIA Quantum-2 InfiniBand networking. Crusoe operates the cluster. The lab pays for the output.

The size of the deal is not the signal. The milestone behind it is. Crusoe Managed Inference crossed $100 million in contracted annual recurring revenue less than a year after launch, and all of it is signed. Training capacity and inference capacity have become separate purchases. Only one of them still requires the buyer to own the stack.

A managed inference endpoint is an odd product when you write it down. You do not get GPUs. You get a service that returns tokens, backed by a cluster the vendor owns, tunes, and patches.

Crusoe calls its version a Tailored Deployment. The customer gets a dedicated, benchmarked, SLA-backed endpoint. Crusoe runs the cluster directly. The customer sends traffic and pays for what it serves. There is no rack to cable and no serving stack to upgrade.

Thinking Machines Lab signed up for that on September 23. The lab will serve production traffic for its Inkling models, GLM 5.2 and 5.3, and its own fine-tuned variants on a dedicated deployment of NVIDIA HGX B200 systems connected with NVIDIA Quantum-2 InfiniBand networking. The agreement runs $65 million a year.

Myle Ott, the lab’s ML infrastructure lead, said the managed service met the standards the lab holds its own systems to. The harder reading is simpler. A lab that prizes infrastructure quality rented the serving layer.

Diagram of how AI procurement splits, showing training capacity financed and owned on one side, inference capacity rented as a managed endpoint on the other, and the Crusoe agreement figures along the bottom.
Training capacity gets financed. Inference capacity gets rented. The line between them is now a buying decision.

The $65 million is the wrong number to watch

Breadth matters as much as any single logo. Crusoe announced a multi-year agreement with Perplexity the week before this deal, then closed a $3.9 billion Series F round. Managed serving is the fastest moving line in a business that started out renting GPU capacity.

The number Crusoe attached to the deal matters more than the deal. Crusoe Managed Inference passed $100 million in contracted annual recurring revenue less than a year from launch.

Contracted is the word doing the work. Not trials, not evaluations, not cloud credits. Signed annual commitments. That is a demand signal for managed serving as a category, not for one lab’s budget.

$65 million is a rounding error in a market where single capacity leases run into gigawatts. As evidence that labs will buy inference instead of building it, it is a different thing entirely.

Inference is becoming its own purchase order

Training capacity arrives through direct chip deals, power contracts, and multi-year commitments to a physical site. Inference capacity needs none of that. The workload is elastic, the pricing is per token, and the switching cost is low.

The buying pattern follows. Labs lock training down with capital. They shop inference among the neoclouds on price, latency, and reliability. Crusoe competes on that list against CoreWeave, Lambda, Together AI, and the hyperscalers’ own serving products. The open serving engines underneath are similar, and vLLM is not the differentiator. Running it well is.

That is the point of a $65 million contract. Running your own inference stack stopped being the default. It became a choice, and an expensive one.

Related reading. Anthropic leased 2.16 gigawatts in Australia for inference alone, and AWS made GPU aware inference routing an EKS add on.

Three questions for your own stack. Do you know what share of your AI bill is training and what share is inference. If your serving team disappeared tomorrow, could you describe the endpoint you would rent. And is the capacity you depend on contracted, or is it best effort.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

One thought on “Thinking Machines Lab Bought Inference Instead of Building It. That Is the New Line Item.”

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.