Security Model and Isolation Strategy for Execute Command Sandboxing in DeepAgents

DeepAgents employs a three-layer security model that combines backend-specific container isolation, base64-encoded JSON payloads to eliminate command injection, and configurable execution timeouts to secure arbitrary shell command execution.

The langchain-ai/deepagents repository implements a robust sandboxing architecture for AI agents that need to execute untrusted shell commands. The security model and isolation strategy employed for the execute command sandboxing ensures that every command runs inside a dedicated, isolated runtime—whether a Runloop devbox, Modal container, or Daytona session—while preventing injection attacks through strict payload encoding and enforcing temporal bounds on execution.

Backend-Specific Runtime Isolation

The foundation of DeepAgents' security model rests on backend-specific isolation, where each concrete sandbox implementation delegates command execution to a separate, container-like environment provided by the underlying service. The host process never receives a live shell; instead, it receives a structured ExecuteResponse containing only stdout, stderr, exit code, and a truncation flag.

Runloop isolates commands by sending them to a Runloop devbox via devbox.cmd.exec, which executes inside an isolated VM managed by the Runloop service.

Modal executes commands inside a Modal Sandbox object using sandbox.exec("bash", "-c", command), providing container-level isolation within Modal's serverless infrastructure.

Daytona creates a fresh session for every command using process.create_session followed by process.execute_session_command, ensuring each command runs within an isolated Daytona sandbox session with no persistent state leakage between invocations.

Payload Encoding and Injection Protection

To prevent command injection and bypass OS argument-size limits (ARG_MAX), the core BaseSandbox class in libs/deepagents/deepagents/backends/sandbox.py constructs all file-operation commands—such as read, write, edit, and glob—using base64-encoded JSON payloads passed through a heredoc (<<'__DEEPAGENTS_EOF__').

This design eliminates the need to interpolate raw user data into shell strings. Instead of concatenating user input directly into a command, the system encodes the operation parameters into JSON, base64-encodes the result, and passes it as a structured payload:


# From BaseSandbox – write command uses base64/JSON payload

payload = json.dumps({"path": file_path, "content": content_b64})
payload_b64 = base64.b64encode(payload.encode("utf-8")).decode("ascii")
cmd = _WRITE_COMMAND_TEMPLATE.format(payload_b64=payload_b64)
result = self.execute(cmd)  # Delegated to concrete backend

By using a heredoc with a quoted delimiter (<<'__DEEPAGENTS_EOF__'), the shell performs no variable expansion or command substitution on the payload content, ensuring that malicious input cannot escape the embedded data context.

Execution Timeout Controls

Every backend supplies a default timeout of 30 minutes (self._default_timeout = 30 * 60). Callers can override this per-invocation; a value of 0 is interpreted as "run forever" (supported in Modal and Daytona implementations).

The timeout is enforced by the underlying service rather than the host process:

  • Runloop: devbox.cmd.exec(..., timeout=…)
  • Modal: sandbox.exec(..., timeout=…)
  • Daytona: process.execute_session_command(..., timeout=…)

This prevents long-running or hung commands from consuming resources indefinitely.

Core Implementation Files

The sandboxing architecture is implemented across the following source files:

Practical Code Examples

Running Commands in a Modal Sandbox

Execute an ls command inside an isolated Modal container:

from deepagents.backends.sandbox import BaseSandbox, ModalSandbox
from deepagents.backends.protocol import ExecuteResponse
import modal

modal_sandbox = modal.Sandbox()
sandbox: BaseSandbox = ModalSandbox(sandbox=modal_sandbox)

resp: ExecuteResponse = sandbox.execute("ls -l /app")
print(resp.output)      # → sandbox-only file list

print(resp.exit_code)   # → 0 on success

Safe File Writing via Base64 Encoding

Create a temporary file without risk of shell injection using BaseSandbox's write() helper:

from deepagents.backends.sandbox import RunloopSandbox

# Assume `devbox` is a Runloop Devbox instance

sandbox = RunloopSandbox(devbox=devbox)

write_res = sandbox.write("/tmp/hello.txt", "Hello, world!")
print(write_res.path)   # → "/tmp/hello.txt"

The write() method internally encodes the content and path as base64 JSON, eliminating injection vectors.

Custom Timeout with Daytona

Run a long-running command with a 2-minute timeout limit:

from deepagents.backends.sandbox import DaytonaSandbox
import daytona

daytona_sandbox = daytona.Sandbox.create()
sandbox = DaytonaSandbox(sandbox=daytona_sandbox, timeout=120)  # 2 min limit

resp = sandbox.execute("sleep 30 && echo done")
print(resp.output)  # → "done"

Summary

  • DeepAgents implements a three-layer security model for execute command sandboxing: backend-specific isolation, payload encoding, and timeout enforcement.
  • Runtime isolation delegates execution to dedicated environments (Runloop VMs, Modal containers, Daytona sessions), ensuring commands never run on the host.
  • Base64-encoded JSON payloads transmitted via heredocs prevent command injection by avoiding shell interpolation of user data.
  • Default 30-minute timeouts protect against resource exhaustion, with per-invocation override capability.
  • The core logic resides in BaseSandbox (libs/deepagents/deepagents/backends/sandbox.py), with backend-specific implementations in the respective partner library directories.

Frequently Asked Questions

How does DeepAgents prevent command injection in sandboxed execute commands?

DeepAgents prevents command injection by encoding all file-operation parameters as base64-encrypted JSON payloads passed through shell heredocs. This method avoids interpolating user data directly into command strings, ensuring that special characters or malicious sequences are treated as literal data rather than executable code. The BaseSandbox class in libs/deepagents/deepagents/backends/sandbox.py handles this encoding automatically for all supported operations.

What is the default timeout for sandboxed command execution?

The default timeout is 30 minutes (1800 seconds), defined as self._default_timeout = 30 * 60 in the BaseSandbox class. This limit is enforced by the underlying backend service (Runloop, Modal, or Daytona), not by the host process. Callers can override this default per-invocation by passing a specific timeout value to the sandbox constructor or execute method.

Which backend isolation methods does DeepAgents support?

DeepAgents supports three backend isolation strategies: Runloop (isolated VMs via devbox.cmd.exec), Modal (serverless containers via sandbox.exec), and Daytona (ephemeral sessions via process.create_session). Each backend implementation resides in its respective partner library directory under libs/partners/ and inherits from the abstract BaseSandbox class to ensure consistent security semantics.

How does the BaseSandbox class handle file operations securely?

The BaseSandbox class handles file operations by constructing commands that embed base64-encoded JSON payloads rather than raw shell arguments. For example, when writing a file, the method serializes the path and content into JSON, encodes it with base64, and injects it into a heredoc template. This approach bypasses shell argument limits (ARG_MAX) and prevents injection attacks by ensuring the shell never interprets user content as commands or variables.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →