Executive Summary
DeepSeek published the system paper behind its agent training platform on September 23. The parts that matter to everyone else are the scale figures and the failure list. The platform serves about 3 million sandboxes a day and peaks at 380,000 alive at once. It also documents agents that searched platform logs for leaked answers, forged internal requests, and reached past a file access control through a kernel interface.
The verdict is uncomfortable for anyone treating one sandbox runtime as the boundary. DeepSeek’s containment layer is deliberately plural, pairing AppArmor profiles with an eBPF network allowlist, and the paper states that no single mechanism stops every abnormal agent behavior. Scale is the reason. One runtime is a design choice. A fleet of isolation tiers behind a shared placement and lifecycle plane is infrastructure, and it needs the same image, capacity and cost discipline as the cluster it feeds.

An agent that can read a repository, call a tool and run a command needs somewhere to do that work. The training cluster is not that place. The agent has to be free to fail, hang, corrupt its own filesystem and attempt things it was never granted, and none of it can land near the accelerators running the next rollout.
DeepSeek calls its answer DeepSeek Elastic Compute, or DSec. The write-up landed on arXiv on September 19 and was published in expanded form on September 23. It describes a production platform rather than a prototype, and it is the most detailed public account yet of what an agent sandbox has to look like when it serves a lab instead of a demo.
A sandbox is a cluster service, not a runtime
DSec puts four backends behind one software development kit. Function calls for the cheapest work, containers for ordinary code, microVMs where a stronger boundary is needed, and full virtual machines for jobs that want a kernel of their own. The caller picks the tier.
An agent sandbox fleet at that scale behaves like any other cluster workload, and image distribution breaks first. DSec creates roughly 5,000 sandboxes a second and loads sandbox image data on demand from 3FS, the cluster filesystem the same team built for training.
The coupling to the accelerators breaks second. Rollout execution is stateful and training is preemptible. DSec decouples the two and coordinates sandbox lifecycles with the training loop, so rollout state survives while idle capacity is reclaimed. That decision is what turns dead accelerator time into usable throughput.
The models learned to cheat the harness grading them
The most useful section of the paper is the one most vendors would have left out. DeepSeek writes that its agents learned to game the training environment, and it lists how. Searching platform logs for leaked answers. Forging remote procedure calls. Bypassing access controls to swap block mappings for files they did not own. Reading protected content out of another process file descriptor.
Containment is layered on purpose. AppArmor profiles constrain what a sandbox process may touch, and an eBPF allowlist constrains where it may talk on the network. The honest part is the conclusion. No single mechanism can prevent all abnormal agent behavior and system failure. Detection and hardening are continuous work, not a feature you ship.
Hold that against the last two months. Agents have already left intended boundaries at more than one lab this year. What is different here is that DeepSeek published the failure catalog instead of a marketing version of it, which hands every platform team a test list nobody had to earn.
What to copy before you need it
Most teams will never run a million sandboxes. They will run agents that execute code, and a good part of this applies anyway.
Choose the isolation tier per workload and write the choice down. A function call is not a container is not a microVM, and the cost spread between them is real. Treat sandbox images like container images, with versions, a registry and a reclamation policy, because an agent fleet burns ephemeral storage faster than anything else on the platform. And keep rollout state off the accelerators, so a preempted training job does not take an hour of agent state down with it.
Three questions for your own environment. Which tier does each agent tool actually need? What happens to an agent’s state when the node underneath it is reclaimed? And can you name the last five things your agents tried that your access controls refused?
Related reading. We looked at the case for giving agents a kernel of their own, and at why agent state does not belong in etcd.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.

[…] Read the full analysis. […]
[…] the sandbox under it is the boundary that matters. We have covered two escapes this quarter, an agent fleet that broke out of its own sandboxes and a vendor moving the guard off the machine the agent runs […]