Photo by Vitaly Gariev on Unsplash. Source: https://unsplash.com/photos/woman-working-late-at-a-computer-in-an-office-5rG70Pk9vjs (Unsplash License).

Executive Summary

Google published benchmark figures for GKE Pod snapshots on September 21, and they are large. Startup latency falls by as much as 89 percent. A 70 billion parameter model comes back in 37 seconds, an 8 billion parameter model in 15. Codeway’s Retake platform went from roughly a minute of cache warming down to eight seconds.

The mechanism is checkpoint and restore, not caching. The feature has been generally available since May on GKE 1.35.3 and later, and it depends on gVisor, so GKE Sandbox must be enabled on the cluster. The verdict for platform teams is that the cold start problem moved rather than closed. It is no longer how fast a model loads but how you manage a snapshot lifecycle and the compatibility hash that decides which pod may restore from which snapshot. Get that wrong and you still pay one cold start per rollout.

Diagram contrasting two startup paths for a Kubernetes inference pod. The cold path runs six phases from scheduling to first token, while the snapshot path restores a frozen pod and streams memory in behind the running process.
Same pod, two paths. The snapshot skips every phase it already paid for once.

Autoscaling an inference service looks free until a new replica comes up. Every replica pulls the container image, downloads the weights, compiles kernels, captures CUDA graphs and profiles the key value cache before it answers a single request. On a model with tens of billions of parameters that sequence is most of the wall clock, and no autoscaler policy shortens it.

GKE Pod snapshots take the other route. Load the model once, freeze the pod, and start every later replica from the frozen state.

Checkpoint and restore is not caching

A snapshot is not a warm image cache. gVisor, the sandboxed runtime underneath, captures the running state of the pod. Memory contents, CPU registers and threads, open file descriptors, GPU memory through NVIDIA checkpoint tooling, the container root filesystem, EmptyDir volumes, tmpfs mounts and listening sockets all go into the snapshot, and the snapshot is stored in Cloud Storage.

Restore is fast partly because it is lazy. Kernel state returns first and the application resumes while memory pages stream in behind it, with page faults fetching whatever is still missing. That is also the trap. A restored pod can report itself healthy and answer a network probe before the model server accepts inference requests, so a readiness probe that only checks whether the process exists will route traffic to a pod that cannot serve.

The published figures are specific. Up to 89 percent lower startup latency, a 70B model restored in 37 seconds, an 8B model in 15. Codeway, running its Retake platform, had built its own cache for compiled artifacts and got startup down to about a minute. It reports eight seconds now, and the team starts H100 instances for a specific job and shuts them down when the job finishes.

The compatibility hash is the part that will bite

Two Kubernetes custom resources carry the configuration. One points at the Cloud Storage bucket. The other selects pods by label, sets the trigger to workload or manual, and sets retention through a last access timeout and a cap on snapshots per group.

Matching a pod to a snapshot is where the operational work sits. GKE hashes the runtime critical fields of the pod, calls the result a distilled pod spec, and embeds it in the snapshot. A restoring pod has to produce an identical hash. The node has to match too, in machine series and CPU architecture, so an N2 snapshot restores on N2 and a G2 on G2.

That turns model digest, CUDA and driver version, GPU type and topology, and runtime configuration into one compatibility key. Change any of them and the snapshot is dead weight. It also means an ordinary rollout still starts one pod cold, the first one, which then becomes the source for the rest. Practitioners are already asking about what the documentation leaves open, including rehydration of secrets, DNS and downstream connections after a restore.

What to check before you switch it on

Pod snapshots need GKE Sandbox, so on Standard clusters you need a node pool with gVisor enabled. Autopilot clusters already have it. Weigh that against your own threat model before adopting the feature for its speed alone.

Decide retention before your first rollout, because snapshots are cheap to create and easy to keep. Bound the number per group and set the last access timeout deliberately rather than by default. Test the readiness probe against a restored pod, not a cold one, since the restored state is the one that serves. And count the cold starts you still pay, one for every changed deployment, because that number is your floor.

Three questions for your own cluster. How much of your scale-up latency is model load rather than scheduling? What in your pod spec changes often enough to invalidate a snapshot? And who owns the bucket those snapshots live in?

Related reading. We broke down why the topology of a GPU cluster decides what the scheduler can do, and why token prices keep moving once caching enters the picture. The full detail is in Google’s Pod snapshot documentation.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.