DeepSeek DSec runs 3 million agent sandboxes a day; on-demand images cut disk writes 57%

DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale

Jialiang Huang, Hongxuan Tang, Jingchang Chen, Yuxuan Liu, Yixiao Chen, Yuan Cheng, Yi Tao, Jingli Zhou, Yupeng Chen, Haoyu Chen, Jiarui Wang, Shengkai Lin, Chuqi Zhang, Bryan Lee Teng, Lian Guo, Zhe Fu, Wenjun Gao, Yisong Wang, Liang Zhao, Zehao Wang, Ziwei Xie, Yongqiang Guo, Peixin Cong, Ziyi Gao, Shuiping Yu, Hanwei Xu, Zuofan Wu, Zhizhou Ren, Yuyang Zhou, Bowei Zhang, Zhihuan Huang, Qihao Zhu, Lei Wang, Tianle Lin, Han Yu, Jiewen Hu, Dejian Yang, Shuo Yang, Shanghao Lu, Shaoyuan Chen, Junjie Qiu, Zhangli Sha, Yinmin Zhong, Yongtong Wu, Shiyu Wang, Wei Liu, Bingzheng Xu, Longhao Chen, Qiushi Du, Yuzhen Huang, Shirong Ma, Yaohui Wang, Mingshu Chen, Tongrui Xiong, Y. C. Yan, Haowen Luo, Haofen Liang, Xiaokang Zhang, Weihao Zeng, Runxin Xu, Peiyi Wang, Jinhua Zhu, Ruoyu Zhang, Wenkai Yang, Qi Tang, Jiping Yu, Tian Ye, Ruizhe Pan, Honghui Ding, Xiaodong Liu, Lingxiao Luo, Zhihong Shao, Yuhan Wu, Jibai Lu, Wen Liu, Haoling Zhang, Jingcheng Hu, Yaoyang Ye, Chaofan Lin, Zhaochen Zhang, Jianan Tong, Hengxu Wu, Zhihao Li, Yicheng Wang, Luyao Wang, Yuzhuo Bai, Lingyue Fu, Ruifan Xu, Y. Z. Wang, Zonglin Li, Mingqi Wei, Haiyang Shen, Chengyuan Zhang, Chao Jin, Zili Zhang, R. H. Yang, Xinbo Xu, Jian Zhou, Ruidong Zhu, Yuzhe Guo, Zelun Pan, Shaoheng Nie, Erhang Li, Shuhan Lin, Zheng Liu, Anshuo Chen, Zilong Lyu, Sinuo Cao, Rui Yu, Chuhao Wang, Junyi Guo, Junxiao Song, Kaifeng Chen, Menghao Ye, Junxian Li, Di Wu, Haiyang Ma, Yilun Wang, Haoran Yang, Yizai Cai, Shichun Liu, Yiping Wang, Junbo Sun, Shicheng Xu, Xiao Bi, Ying He, Yichao Zhang, Mingxing Zhang, Liyue Zhang, Panpan Huang, Wenfeng Liang

cs.DC

2026-09-19

DeepSeek's DSec is an elastic sandbox platform for agent RL: ~160 nodes, 3 million instances a day, 380K peak concurrency. On-demand image loading finishes an 8,192-container burst in 35 minutes versus 60 for eager pull, with 57% fewer disk writes.

What problem this solves

Agentic training puts a model in a real environment: it walks a repo, runs commands, opens a browser, edits files, and takes exit codes or test pass rates as reward. Those environments have to be isolated, stateful, and long-lived. A single job can ask for 32K sandboxes. Image diversity is high and reuse is low, so a node-local cache cannot absorb the working set. While the sandbox waits for the next model action, CPU sits idle and memory plus writable state stay pinned. GPU trainers get preempted; rollouts cannot die with them.

Wrapping a container runtime is not enough. DeepSeek built DSec as a production elastic platform and ran it from V3.2 through V4.1.

Method

Callers use the Python SDK libdsec and pick a backend. FnCall covers short stateless work (OJ tasks, compilation, GPU kernels) inside precreated containers. Containers are the default for software engineering and tool use: fast start, high density, shared host kernel. Firecracker microVMs buy a stronger isolation boundary at higher memory and startup cost. Full VMs via QEMU cover Android, GUIs and rendering. Containers and microVMs dominate production count and resources.

At cluster level, IAM authenticates, the apiserver is a stateless ingress, placement filters healthy nodes then samples a few and picks the least loaded, and a watcher probes liveness. The per-node edge does local admission. Container and VM sandboxes run an aether proxy plus chronus shell sessions; FnCall skips that path. FnCall and containers sit inside QEMU/libvirt VMs, adding a kernel and network boundary.

Environments are independently versioned layers: OS base, task workspace, toolkits such as DeepSeek Harness. dockerd is patched to compose overlayfs at create time. Immutable layers are EROFS, compressed and randomly readable. MicroVMs expose EROFS as read-only block devices and overlay inside the guest. One production week: 11,266 container base images and 102,171 workspaces (82.8 TB); two microVM bases and 53,590 workspaces (50.9 TB). 103 toolkits; 67.8% of sandboxes need at least one extra workspace or toolkit layer.

Images live on 3FS. Containers keep metadata local and pull file data on access; writes stay on local disk. MicroVM writable disks use OverlayBD plus ublk with 256 KiB chunks and a second-level local cache. Sampled images touch only 4.2% to 13.3% of their bytes at runtime.

Density comes from overcommit. About 90% of sandboxes average at most 5% of requested CPU. Production has held 3,200 containers or 800 microVMs per node. MicroVMs share the host page cache for read-only layers via virtio-pmem with DAX, and return cold pages with DAMON plus virtio-balloon free-page reporting. Latency-sensitive work gets Linux core scheduling; best-effort work runs SCHEDIDLE so SMT siblings do not fight.

Co-design with the RL stack is explicit. From V4.1 the agent loop leaves the preemptible GPU pool; a worker container plus agent sandbox become the source of truth for rollout state. On preemption, containers docker pause then memory.reclaim; microVMs snapshot and kill Firecracker. Agents build environments on the same platform; packdiff takes incremental snapshots. Per-task eBPF allowlists and AppArmor on files and sockets are there to stop answer-seeking through logs, internal sockets and package mirrors.

When on-prem utilization exceeds 80%, tasks whose images sit in a 30 TB shared set spill to cloud VMs. 200 cloud VMs absorb about 30% of peak overflow.

Results

One scale unit is about 160 CPU nodes, 30K cores and 250 TB of DRAM. It serves about 3 million sandboxes a day, peaks around 380K concurrent, and creates more than 5,000 sandboxes per second. Median lifetime is 17.4 minutes for containers and 15.5 for microVMs; p99 exceeds three hours for both.

Evaluation used a separate 10-node cluster on internal software-engineering work, SWE-bench, Terminal-Bench and security tasks. An 8,192-container burst: on-demand EROFS finishes in about 35 minutes, matching the fully local baseline; cold Docker pull takes over 60 minutes, 1.71× slower. Cumulative disk writes per node drop from over 1,600 GB to about 700 GB, 57% less, close to the 600 GB local baseline.

Provisioning workspaces as EROFS mounts versus tar.gz extraction: 45 minutes versus 79 minutes end to end, 1.76× faster. The tar path writes about 5.5× more bytes and 3.4× higher peak throughput.

Memory: virtio-pmem cuts peak host use 40.2% versus baseline; DAMON plus free-page reporting cuts time-integrated memory 21.2%; both together are lowest. pmem lifts transient CPU from 26.5% to 41.4%, so a CPU-tight fleet can keep virtio-blk and enable reclamation only.

CPU: at 50% best-effort load, an unprotected latency-sensitive chess agent slows 45.2% per step. SCHEDIDLE alone recovers at most 3.4%. Adding core scheduling caps inflation at 17.3%. Residual delay is turbo drop plus LLC and memory-bandwidth contention; they did not add bandwidth isolation.

Why it matters

The bottleneck in agent RL is often how many real environments can stay alive, and whether image bring-up stalls the training loop, not a missing GPU kernel. DSec turns that into a cluster product that overcommits, survives GPU preemption, and applies per-task network policy. Anyone using inference sandboxes such as E2B or Code Interpreter gets the training-scale numbers: 32K instances per job, low image fanout, state that must outlive preemption.

The storage bet is concrete. 3FS is weak on small random I/O, so writes stay local, reads are on demand, and metadata prefers the node. This is a platform report, not a new isolation primitive.

Limitations

The evaluation covers the §5 infrastructure mechanisms. Preemption recovery and reward-hacking defenses in §6 have no controlled numbers. The testbed is 10 nodes, not the 160-node production unit. Latency numbers come from a chess agent; GUI or compile-heavy tasks need their own measurement.

Access control does not stop kernel bugs. An agent that grepped from / into /proc crashed the kernel. Another used XFSIOCSWAPEXT to dodge file protections and took the filesystem down. When chronus records stdout for async retrieval, yes wrote tens of gigabytes. Hardening is ongoing, not a one-shot policy.

There is no cost model and no head-to-head against E2B or RunD on the same load. Cloud bursting only accepts container tasks whose images sit in that 30 TB set.

Terms

Source

What people are saying

Related papers

All paper explainers