Security Considerations for Coding Agents: Sandbox and Command Parsing Explained

Coding agents that execute model-generated code must run inside hardened Docker sandboxes with network isolation, capability-dropped containers, and strict resource limits to prevent privilege escalation, denial-of-service attacks, and data exfiltration.

This guide examines the defense-in-depth security architecture implemented in bojieli/ai-agent-book (Chapter 8), where a self-modifying coding agent evaluates retry-policy code submitted by a language model. The system combines container hardening, deterministic image construction, and runtime validation to safely parse and execute untrusted code.

Sandbox Architecture Overview

The repository's candidate_sandbox.py orchestrates sandbox execution while sandbox_runner.py handles validation inside the container. Together they enforce multiple security layers before any model-generated code executes on the host.

The sandbox design assumes all input is potentially malicious. Every component—including image building, command parsing, and runtime verification—includes hardening against injection, escape, and resource exhaustion attacks.

Container Isolation and Hardening

Network and Privilege Restrictions

The Docker container launches with aggressive lockdown flags that eliminate common escape vectors. In candidate_sandbox.py lines 74–90, the _docker_command function constructs a container with:

  • --network none — completely disables network access
  • --cap-drop ALL — removes all Linux capabilities
  • --security-opt no-new-privileges:true — prevents privilege escalation via setuid binaries
  • --read-only — mounts the root filesystem as read-only
  • Non-root user 65534:65534 (the nobody user)

These flags prevent the candidate code from accessing host network interfaces, escalating to root, or modifying system files.

Resource Limiting

Hard ceilings stop resource exhaustion attacks. The same _docker_command specifies:

"--memory", "64m",           # 64 MB RAM limit

"--memory-swap", "64m",      # No swap allowance

"--cpus", "0.5",             # Half a CPU core

"--pids-limit", "16",        # Maximum 16 processes

"--ulimit", "cpu=2:2",       # 2-second CPU time limit

"--tmpfs", "/tmp:rw,noexec,nosuid,nodev,size=16m,mode=1777",

The minimal 16 MB tmpfs at /tmp provides temporary workspace without executable permissions or setuid capabilities.

Deterministic and Verifiable Image Construction

Content-Addressed Image Tags

Supply-chain attacks are mitigated through cryptographic verification of the sandbox image. The _default_image function in candidate_sandbox.py (lines 29–32) computes:

def _default_image() -> str:
    # Hash of Dockerfile + runner ensures reproducibility

    digest = hashlib.sha256(DOCKERFILE.read_bytes() + RUNNER.read_bytes()).hexdigest()[:12]
    return f"ai-agent-book/self-modifying-agent-sandbox:{digest}"

This content-addressed tagging guarantees that any image modification—even a single byte change in the Dockerfile or runner script—produces a distinct tag. Operators can audit exactly what code runs in production.

Controlled Image Sources

Environment variable SELF_MODIFY_SANDBOX_IMAGE allows deployment-specific images, but only when explicitly configured (lines 55–66). Without this override, the system builds fresh from the controlled Dockerfile.sandbox, preventing silent substitution of compromised base images.

Input Validation and Command Parsing Security

Size Limits Prevent Payload Attacks

candidate_sandbox.py enforces strict bounds (lines 20–23, 99–106):

  • MAX_REQUEST_BYTES = 1_048_576 (1 MiB) — rejects oversized HTTP payloads
  • MAX_SOURCE_BYTES = 262_144 (256 KiB) — limits individual source files
  • MAX_OUTPUT_BYTES = 1_048_576 (1 MiB) — caps stdout/stderr capture

These limits stop memory exhaustion attacks through maliciously large inputs or log-bombing via excessive output.

Strict JSON Parsing

The request parser validates structure before processing (lines 35–44, 141–144):

def _parse_response(raw: bytes) -> dict:
    if len(raw) > MAX_OUTPUT_BYTES:
        raise SandboxError(f"Output exceeds {MAX_OUTPUT_BYTES} bytes")
    try:
        return json.loads(raw.decode("utf-8"))
    except (json.JSONDecodeError, UnicodeDecodeError) as e:
        raise SandboxError(f"Invalid JSON response: {e}")

Malformed payloads, encoding attacks, or oversized responses trigger immediate SandboxError exceptions without passing data to downstream processing.

Runtime Validation Inside the Sandbox

Static Analysis of Candidate Code

sandbox_runner.py performs multi-layer verification before executing any policy logic (lines 45–62, 66–82):

  1. AST parsing — verifies Python syntax without execution
  2. Required function detection — confirms should_retry and should_open_circuit exist
  3. Signature matching — validates exact function signatures against the public API:
checks["public_api_compatible"] = (
    str(inspect.signature(namespace["should_retry"])) == "(error_code, retryable, attempt)"
    and str(inspect.signature(namespace["should_open_circuit"]))
    == "(consecutive_failures, *, error_code='', retryable=True)"
)
  1. Behavioral testing — replays failure trajectories to verify correct circuit-opening and temporary-recovery logic

This static-then-dynamic approach catches malicious or malformed code before it executes outside controlled test scenarios.

Error Handling and Containment

Subprocess Timeouts and Cleanup

All container interactions include timeout guards (lines 77–86, 188–196):

try:
    result = subprocess.run(
        cmd,
        input=request_json,
        capture_output=True,
        timeout=timeout_seconds or DEFAULT_TIMEOUT_SECONDS,
        check=True,
    )
except subprocess.TimeoutExpired:
    _kill_container(name)
    raise SandboxError(f"Sandbox timed out after {timeout_seconds}s")

Runaway processes are terminated via _kill_container, and diagnostic information is deliberately limited to prevent information leakage to attackers.

Bounded Output Capture

Stdout and stderr are captured in fixed-size byte arrays. Overflow beyond MAX_OUTPUT_BYTES is treated as a sandbox error rather than truncation, preventing attackers from hiding malicious output in excessive noise.

Defense-in-Depth Summary

Layer Mechanism Implementation Location
Image integrity SHA-256 content hashing candidate_sandbox.py _default_image
Build control Dockerfile-only builds without override candidate_sandbox.py lines 55–66
Network isolation --network none, --ipc none candidate_sandbox.py _docker_command
Privilege restriction Capability drop, no-new-privileges, non-root user candidate_sandbox.py lines 74–90
Resource containment Memory, CPU, PID, and filesystem limits candidate_sandbox.py _docker_command
Input validation Size limits, strict JSON parsing candidate_sandbox.py lines 20–23, 35–44
Code verification AST analysis, signature matching, behavioral tests sandbox_runner.py lines 45–82
Execution monitoring Timeouts, bounded output, error containment candidate_sandbox.py lines 99–106, 188–196

Summary

  • Docker hardening with --network none, --cap-drop ALL, and read-only filesystems eliminates common container escape vectors for coding agents.
  • Content-addressed image tags in candidate_sandbox.py ensure only explicitly verified sandbox code executes, preventing supply-chain compromises.
  • Strict input limits (1 MiB requests, 256 KiB sources) and JSON validation block memory exhaustion and injection attacks at the parsing layer.
  • Multi-stage validation in sandbox_runner.py combines static AST analysis, signature verification, and dynamic behavioral testing before any policy code runs against real data.
  • Comprehensive timeouts and bounded I/O guarantee that malicious or buggy code cannot hang the host or exfiltrate data through excessive output.

Frequently Asked Questions

How does the sandbox prevent network-based attacks?

The Docker command explicitly passes --network none and --ipc none, disabling all network interfaces and inter-process communication. The sandbox runner contains no networking imports, and the content-addressed image build ensures no unauthorized network tools are present. Even if an attacker injected network code, the container has no route to external services.

What happens if model-generated code tries to exhaust CPU or memory?

Hard limits terminate the attack: 64 MB RAM with no swap, 0.5 CPU cores, 16 process slots, and 2-second CPU time via ulimit. The subprocess.run timeout aborts runaway execution, and _kill_container forcibly removes the Docker container. Excessive output beyond 1 MiB triggers a SandboxError rather than buffer growth.

Why verify function signatures instead of just running the code?

Static signature verification in sandbox_runner.py catches interface violations before behavioral testing begins. This fails fast on malformed or malicious submissions—such as functions with unexpected parameters that might exploit test harness internals—without exposing the full evaluation logic to arbitrary code execution.

Can the sandbox image be replaced with a malicious version?

Only if the SELF_MODIFY_SANDBOX_IMAGE environment variable is explicitly set by administrators. Without this override, candidate_sandbox.py rebuilds from the controlled Dockerfile.sandbox on every invocation. Operators can audit the DOCKERFILE and RUNNER hashes against expected values to detect tampering.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →