How Execution and Inactivity Timeouts Are Handled in Background Agents

Background agents use configurable timeouts managed by the SandboxSupervisor in entrypoint.py, with execution timeouts enforced via asyncio.wait_for and inactivity timeouts monitored through deadline checks in the SSE bridge.

The ColeMurray/background-agents repository implements a robust timeout system to prevent runaway processes in sandboxed environments. The Python SandboxSupervisor orchestrates script execution and network connectivity through multiple configurable timeout mechanisms that protect against hanging hooks and stalled connections.

Execution Timeouts for Setup and Start Scripts

The SandboxSupervisor enforces strict time limits on user-defined lifecycle scripts using asyncio.wait_for to ensure sandbox initialization cannot hang indefinitely.

Default Values and Configuration

Two primary execution timeouts protect the setup and start phases:

  • DEFAULT_SETUP_TIMEOUT_SECONDS defaults to 300 seconds (5 minutes) for setup.sh scripts
  • DEFAULT_START_TIMEOUT_SECONDS defaults to 120 seconds (2 minutes) for start.sh scripts

You can override these defaults by setting environment variables before the supervisor initializes:

export SETUP_TIMEOUT_SECONDS=45   # Reduce from 300s to 45s

export START_TIMEOUT_SECONDS=60   # Reduce from 120s to 60s

These values are resolved through the _resolve_timeout_seconds helper method in packages/sandbox-runtime/src/sandbox_runtime/entrypoint.py around lines 1520–1540.

Implementation in _run_hook

The actual enforcement occurs in the _run_hook method (lines 1500–1600 in entrypoint.py). The supervisor spawns a subprocess and wraps communication in a timeout:


# Inside SandboxSupervisor._run_hook (excerpt)

timeout_seconds = self._resolve_timeout_seconds(
    timeout_env_var, 
    default_timeout_seconds
)
proc = await asyncio.create_subprocess_exec(...)
stdout, _ = await asyncio.wait_for(
    proc.communicate(), 
    timeout=timeout_seconds
)

If the hook exceeds the configured duration, asyncio.wait_for raises a TimeoutError, triggering the supervisor to kill the subprocess and mark the hook as failed.

SSE Inactivity Timeout for Control-Plane Communication

The SSE bridge maintains a persistent connection to the control-plane, requiring its own timeout mechanism to detect stalled streams.

Bridge Implementation Details

In packages/sandbox-runtime/src/sandbox_runtime/bridge.py (line 247), the bridge initializes the inactivity timeout:

self.sse_inactivity_timeout = self._resolve_timeout_seconds(
    "SSE_INACTIVITY_TIMEOUT_SECONDS", 
    default=120
)

The default 120 seconds matches the Open-Code client’s built-in timeout, ensuring compatibility while preventing indefinite waiting.

Deadline-Based Monitoring

The bridge reads SSE events in an asynchronous loop, resetting a deadline after each successful chunk reception. Around lines 1150–1160 in bridge.py, the logic checks:


# In the SSE read loop:

deadline = loop.time() + self.sse_inactivity_timeout

# ... after receiving data, deadline resets

if loop.time() > deadline:
    self.log.warn(
        "bridge.sse_inactivity_timeout", 
        timeout=self.sse_inactivity_timeout
    )
    break  # abort the stream

To configure a shorter timeout, set the environment variable before launching:

export SSE_INACTIVITY_TIMEOUT_SECONDS=30   # Shorten from 120s to 30s

Additional Timeout Protections

The supervisor implements specialized timeouts for specific operational concerns beyond basic script execution.

Port Readiness Timeout

When waiting for sidecar services to bind, the supervisor uses SIDECAR_TIMEOUT_SECONDS (default 5 seconds). The _wait_for_port method (lines 1145–1159 in entrypoint.py) repeatedly attempts connections until the deadline, logging a "port_readiness.timeout" warning if the port fails to open.

Override this value for slower-starting dependencies:

export SIDECAR_TIMEOUT_SECONDS=10   # Increase from 5s to 10s

MCP Package Installation Timeout

For installing npm packages required by local MCP servers, the supervisor enforces MCP_PACKAGE_INSTALL_TIMEOUT_SECONDS (default 180 seconds). The _install_mcp_packages method wraps the npm install command:

await asyncio.wait_for(
    proc.communicate(),
    timeout=self.MCP_PACKAGE_INSTALL_TIMEOUT_SECONDS
)

On timeout, a "mcp.packages_install_timeout" warning is emitted. Extend this for large dependency trees:

export MCP_PACKAGE_INSTALL_TIMEOUT_SECONDS=300   # 5 minutes

Environment-Based Configuration

All timeout values resolve through a common helper method in entrypoint.py (lines 1520–1540):

def _resolve_timeout_seconds(self, env_var: str, default: int) -> int:
    """Read an integer from the environment, falling back to `default`."""
    raw = os.environ.get(env_var)
    if raw is None:
        return default
    try:
        timeout = int(raw)
    except ValueError:
        return default
    return timeout if timeout > 0 else default

This implementation ensures that:

  1. Environment variables take precedence over hardcoded defaults
  2. Invalid values (non-integers or negative numbers) fall back safely to defaults
  3. Zero or negative timeouts are treated as invalid, preventing accidental disabling

Summary

  • Execution timeouts protect setup.sh and start.sh scripts with 300s and 120s defaults, enforced via asyncio.wait_for in the _run_hook method
  • Inactivity timeouts prevent stalled SSE connections using a 120s deadline-based check in bridge.py
  • Sidecar timeouts wait 5 seconds for port availability, while MCP install timeouts allow 180 seconds for npm operations
  • All values are configurable via environment variables (SETUP_TIMEOUT_SECONDS, SSE_INACTIVITY_TIMEOUT_SECONDS, etc.) parsed through _resolve_timeout_seconds
  • The SandboxSupervisor in entrypoint.py serves as the central authority for timeout orchestration across the sandbox runtime

Frequently Asked Questions

How do I change the default timeout for setup scripts?

Set the SETUP_TIMEOUT_SECONDS environment variable in your sandbox environment. The SandboxSupervisor reads this value during initialization in packages/sandbox-runtime/src/sandbox_runtime/entrypoint.py and applies it to all setup.sh executions via the _run_hook method.

What happens when a background agent hook exceeds its execution timeout?

When a hook exceeds the configured limit, asyncio.wait_for raises a TimeoutError, causing the supervisor to terminate the subprocess and mark the hook as failed. This prevents runaway scripts from consuming resources indefinitely.

Can I disable the SSE inactivity timeout?

No, the timeout cannot be fully disabled through configuration. The _resolve_timeout_seconds method treats zero or negative values as invalid and falls back to the default (120 seconds). You can only adjust the duration by setting SSE_INACTIVITY_TIMEOUT_SECONDS to a positive integer value.

Where is the timeout configuration logic implemented?

The configuration logic resides in packages/sandbox-runtime/src/sandbox_runtime/entrypoint.py within the SandboxSupervisor class. The _resolve_timeout_seconds method (lines 1520–1540) handles parsing for all timeout types, while specific enforcement occurs in _run_hook for execution timeouts and bridge.py for SSE inactivity monitoring.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →