Security Considerations for Coding Agents: Sandbox and Command Parsing Explained
Coding agents that execute model-generated code must run inside hardened Docker sandboxes with network isolation, capability-dropped containers, and strict resource limits to prevent privilege escalation, denial-of-service attacks, and data exfiltration.
This guide examines the defense-in-depth security architecture implemented in bojieli/ai-agent-book (Chapter 8), where a self-modifying coding agent evaluates retry-policy code submitted by a language model. The system combines container hardening, deterministic image construction, and runtime validation to safely parse and execute untrusted code.
Sandbox Architecture Overview
The repository's candidate_sandbox.py orchestrates sandbox execution while sandbox_runner.py handles validation inside the container. Together they enforce multiple security layers before any model-generated code executes on the host.
The sandbox design assumes all input is potentially malicious. Every component—including image building, command parsing, and runtime verification—includes hardening against injection, escape, and resource exhaustion attacks.
Container Isolation and Hardening
Network and Privilege Restrictions
The Docker container launches with aggressive lockdown flags that eliminate common escape vectors. In candidate_sandbox.py lines 74–90, the _docker_command function constructs a container with:
--network none— completely disables network access--cap-drop ALL— removes all Linux capabilities--security-opt no-new-privileges:true— prevents privilege escalation via setuid binaries--read-only— mounts the root filesystem as read-only- Non-root user
65534:65534(thenobodyuser)
These flags prevent the candidate code from accessing host network interfaces, escalating to root, or modifying system files.
Resource Limiting
Hard ceilings stop resource exhaustion attacks. The same _docker_command specifies:
"--memory", "64m", # 64 MB RAM limit
"--memory-swap", "64m", # No swap allowance
"--cpus", "0.5", # Half a CPU core
"--pids-limit", "16", # Maximum 16 processes
"--ulimit", "cpu=2:2", # 2-second CPU time limit
"--tmpfs", "/tmp:rw,noexec,nosuid,nodev,size=16m,mode=1777",
The minimal 16 MB tmpfs at /tmp provides temporary workspace without executable permissions or setuid capabilities.
Deterministic and Verifiable Image Construction
Content-Addressed Image Tags
Supply-chain attacks are mitigated through cryptographic verification of the sandbox image. The _default_image function in candidate_sandbox.py (lines 29–32) computes:
def _default_image() -> str:
# Hash of Dockerfile + runner ensures reproducibility
digest = hashlib.sha256(DOCKERFILE.read_bytes() + RUNNER.read_bytes()).hexdigest()[:12]
return f"ai-agent-book/self-modifying-agent-sandbox:{digest}"
This content-addressed tagging guarantees that any image modification—even a single byte change in the Dockerfile or runner script—produces a distinct tag. Operators can audit exactly what code runs in production.
Controlled Image Sources
Environment variable SELF_MODIFY_SANDBOX_IMAGE allows deployment-specific images, but only when explicitly configured (lines 55–66). Without this override, the system builds fresh from the controlled Dockerfile.sandbox, preventing silent substitution of compromised base images.
Input Validation and Command Parsing Security
Size Limits Prevent Payload Attacks
candidate_sandbox.py enforces strict bounds (lines 20–23, 99–106):
MAX_REQUEST_BYTES = 1_048_576(1 MiB) — rejects oversized HTTP payloadsMAX_SOURCE_BYTES = 262_144(256 KiB) — limits individual source filesMAX_OUTPUT_BYTES = 1_048_576(1 MiB) — caps stdout/stderr capture
These limits stop memory exhaustion attacks through maliciously large inputs or log-bombing via excessive output.
Strict JSON Parsing
The request parser validates structure before processing (lines 35–44, 141–144):
def _parse_response(raw: bytes) -> dict:
if len(raw) > MAX_OUTPUT_BYTES:
raise SandboxError(f"Output exceeds {MAX_OUTPUT_BYTES} bytes")
try:
return json.loads(raw.decode("utf-8"))
except (json.JSONDecodeError, UnicodeDecodeError) as e:
raise SandboxError(f"Invalid JSON response: {e}")
Malformed payloads, encoding attacks, or oversized responses trigger immediate SandboxError exceptions without passing data to downstream processing.
Runtime Validation Inside the Sandbox
Static Analysis of Candidate Code
sandbox_runner.py performs multi-layer verification before executing any policy logic (lines 45–62, 66–82):
- AST parsing — verifies Python syntax without execution
- Required function detection — confirms
should_retryandshould_open_circuitexist - Signature matching — validates exact function signatures against the public API:
checks["public_api_compatible"] = (
str(inspect.signature(namespace["should_retry"])) == "(error_code, retryable, attempt)"
and str(inspect.signature(namespace["should_open_circuit"]))
== "(consecutive_failures, *, error_code='', retryable=True)"
)
- Behavioral testing — replays failure trajectories to verify correct circuit-opening and temporary-recovery logic
This static-then-dynamic approach catches malicious or malformed code before it executes outside controlled test scenarios.
Error Handling and Containment
Subprocess Timeouts and Cleanup
All container interactions include timeout guards (lines 77–86, 188–196):
try:
result = subprocess.run(
cmd,
input=request_json,
capture_output=True,
timeout=timeout_seconds or DEFAULT_TIMEOUT_SECONDS,
check=True,
)
except subprocess.TimeoutExpired:
_kill_container(name)
raise SandboxError(f"Sandbox timed out after {timeout_seconds}s")
Runaway processes are terminated via _kill_container, and diagnostic information is deliberately limited to prevent information leakage to attackers.
Bounded Output Capture
Stdout and stderr are captured in fixed-size byte arrays. Overflow beyond MAX_OUTPUT_BYTES is treated as a sandbox error rather than truncation, preventing attackers from hiding malicious output in excessive noise.
Defense-in-Depth Summary
| Layer | Mechanism | Implementation Location |
|---|---|---|
| Image integrity | SHA-256 content hashing | candidate_sandbox.py _default_image |
| Build control | Dockerfile-only builds without override | candidate_sandbox.py lines 55–66 |
| Network isolation | --network none, --ipc none |
candidate_sandbox.py _docker_command |
| Privilege restriction | Capability drop, no-new-privileges, non-root user | candidate_sandbox.py lines 74–90 |
| Resource containment | Memory, CPU, PID, and filesystem limits | candidate_sandbox.py _docker_command |
| Input validation | Size limits, strict JSON parsing | candidate_sandbox.py lines 20–23, 35–44 |
| Code verification | AST analysis, signature matching, behavioral tests | sandbox_runner.py lines 45–82 |
| Execution monitoring | Timeouts, bounded output, error containment | candidate_sandbox.py lines 99–106, 188–196 |
Summary
- Docker hardening with
--network none,--cap-drop ALL, and read-only filesystems eliminates common container escape vectors for coding agents. - Content-addressed image tags in
candidate_sandbox.pyensure only explicitly verified sandbox code executes, preventing supply-chain compromises. - Strict input limits (1 MiB requests, 256 KiB sources) and JSON validation block memory exhaustion and injection attacks at the parsing layer.
- Multi-stage validation in
sandbox_runner.pycombines static AST analysis, signature verification, and dynamic behavioral testing before any policy code runs against real data. - Comprehensive timeouts and bounded I/O guarantee that malicious or buggy code cannot hang the host or exfiltrate data through excessive output.
Frequently Asked Questions
How does the sandbox prevent network-based attacks?
The Docker command explicitly passes --network none and --ipc none, disabling all network interfaces and inter-process communication. The sandbox runner contains no networking imports, and the content-addressed image build ensures no unauthorized network tools are present. Even if an attacker injected network code, the container has no route to external services.
What happens if model-generated code tries to exhaust CPU or memory?
Hard limits terminate the attack: 64 MB RAM with no swap, 0.5 CPU cores, 16 process slots, and 2-second CPU time via ulimit. The subprocess.run timeout aborts runaway execution, and _kill_container forcibly removes the Docker container. Excessive output beyond 1 MiB triggers a SandboxError rather than buffer growth.
Why verify function signatures instead of just running the code?
Static signature verification in sandbox_runner.py catches interface violations before behavioral testing begins. This fails fast on malformed or malicious submissions—such as functions with unexpected parameters that might exploit test harness internals—without exposing the full evaluation logic to arbitrary code execution.
Can the sandbox image be replaced with a malicious version?
Only if the SELF_MODIFY_SANDBOX_IMAGE environment variable is explicitly set by administrators. Without this override, candidate_sandbox.py rebuilds from the controlled Dockerfile.sandbox on every invocation. Operators can audit the DOCKERFILE and RUNNER hashes against expected values to detect tampering.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →