How to Debug Open Agents Execution and Tool Call Issues: A Complete Guide

To debug Open Agents execution and tool call issues, inspect the sandbox initialization in utils.ts, validate working directory resolution in bash.ts, and trace the standardized execution result payload returned by the sandbox interface.

Open Agents by Vercel Labs executes every user request inside an isolated sandbox and routes actions through strictly typed tools. When a tool call fails, times out, or returns unexpected output, the debugging trace spans three distinct layers: the tool definition, the sandbox interface, and the runtime orchestration. This guide walks through the exact source code locations and instrumentation strategies needed to diagnose execution failures in the vercel-labs/open-agents repository.

Understanding the Open Agents Execution Architecture

Open Agents splits execution responsibility across three layers. Debugging requires checking each layer sequentially.

Tool Definition Layer

Typed tools reside in packages/agent/tools/*.ts and handle input validation, approval decisions, and sandbox invocation. For example, bashTool in packages/agent/tools/bash.ts executes shell commands, while grepTool and globTool handle file-system queries. Each tool imports getSandbox from utils.ts to access the isolated execution environment.

Sandbox Interface Layer

The sandbox interface is implemented by @open-harness/sandbox and exposed through helper functions in packages/agent/tools/utils.ts. All sandbox calls return a uniform result object containing success, exitCode, stdout, stderr, and an optional truncated flag. The getSandbox function (lines 71-86) validates that the sandbox exists in the experimental context and throws a descriptive error if initialization failed.

System Prompt and Runtime Orchestration

The core prompt logic in packages/agent/system-prompt.ts builds the agent context and injects model-specific overlays. This file also accepts customInstructions that can enable verbose logging, which surfaces debug information in the execution trace.

Debugging Sandbox Initialization Failures

The most common root cause of execution failures is a missing or uninitialized sandbox. The getSandbox function in packages/agent/tools/utils.ts throws an explicit error when the experimental context lacks a sandbox instance.

Wrap your tool execution in a try-catch block to surface the initialization error:

import { getSandbox } from "./utils";

export const bashTool = tool({
  // ...
  execute: async ({ command, cwd }, { experimental_context, abortSignal }) => {
    let sandbox;
    try {
      sandbox = await getSandbox(experimental_context, "bash");
    } catch (e) {
      // Surface a clear error for the UI
      return {
        success: false,
        exitCode: null,
        stdout: "",
        stderr: (e as Error).message,
      };
    }

    // Proceed with execution...
    const workingDir = cwd
      ? path.resolve(sandbox.workingDirectory, cwd)
      : sandbox.workingDirectory;
    
    const result = await sandbox.exec(command, workingDir, TIMEOUT_MS, {
      signal: abortSignal,
    });

    return {
      success: result.success,
      exitCode: result.exitCode,
      stdout: result.stdout,
      stderr: result.stderr,
      ...(result.truncated && { truncated: true }),
    };
  },
});

If you see the error "Sandbox not initialized in context", verify that the agent's prepareCall function sets experimental_context: { sandbox, ... } before invoking the tool.

Resolving Working Directory and Path Issues

bashTool resolves the optional cwd parameter against the sandbox's workingDirectory (lines 8-14 in packages/agent/tools/bash.ts). A common bug is passing a relative path that resolves outside the sandbox root, which triggers a security error from isPathWithinDirectory (lines 23-33 in packages/agent/tools/utils.ts).

Log the resolved path before execution to catch traversal errors:

const workingDir = cwd
  ? path.isAbsolute(cwd) ? cwd : path.resolve(workingDirectory, cwd)
  : workingDirectory;

console.log("[bashTool] resolved workingDir:", workingDir);

If the sandbox rejects the path, check that cwd is relative to the current task directory and not an absolute system path.

Inspecting Tool Execution Results

All sandbox executions return a standardized result object. When success is false and stderr is empty, the command likely timed out. When truncated is true, the output exceeded the ~50,000 character limit and was cut short.

Check the result shape in packages/agent/tools/bash.ts (lines 45-56):

return {
  success: result.success,
  exitCode: result.exitCode,
  stdout: result.stdout,
  stderr: result.stderr,
  ...(result.truncated && { truncated: true }),
};

Debug tips:

  • Timeout: Increase TIMEOUT_MS or optimize the command (e.g., git log --oneline -n 20 instead of full history).
  • Truncation: Pipe output through head or use more specific search patterns to reduce verbosity.

Handling Dangerous Command Approvals

bashTool runs commandNeedsApproval (lines 31-47 in packages/agent/tools/bash.ts) to check against DANGEROUS_COMMAND_PATTERNS. If a pattern matches, the tool's needsApproval handler consults the optional options.needsApproval callback (lines 52-58).

When a command is rejected, the UI receives a needsApproval boolean. Investigate the caller that supplied options.needsApproval—usually a sub-agent or the parent agent implementing the approval policy.

Override the approval callback for granular control:

import { bashTool } from "packages/agent/tools/bash";

const safeBash = bashTool({
  needsApproval: (args) => {
    // Allow `rm -rf` only inside a temporary directory
    if (/rm\s+-rf/.test(args.command)) {
      return !args.cwd?.startsWith("tmp/");
    }
    return false; // other commands are safe
  },
});

Debugging Sub-Agent Tasks

taskTool in packages/agent/tools/task.ts streams intermediate states back to the UI, including tool-call counts, pending calls, and token usage. The first yielded object contains a stable startedAt timestamp. Later yields contain pending with the sub-agent's internal tool name and raw input (line 21).

Instrument your UI handler to log each yield:

// Inside a UI component that receives taskTool streams
function handleTaskPart(part: TaskToolUIPart) {
  if (part.pending) {
    console.log(
      "Sub-agent waiting on tool:",
      part.pending.name,
      "input:",
      part.pending.input,
    );
  }
  if (part.final) {
    console.log("Sub-agent finished. Message:", part.final);
  }
}

Debug tips:

  • If yields show no finish-step events, the sub-agent is stuck in a loop.
  • If pending never becomes undefined, the sub-agent is waiting for a tool that requires approval.

Enabling Verbose Logging in the System Prompt

The core system prompt in packages/agent/system-prompt.ts accepts customInstructions (lines 28-33) that inject additional directives into the agent context. Add a debug instruction to force the agent to prepend diagnostic information to every tool payload.

Enable debug logging:

customInstructions: "Log every tool call payload and sandbox result."

This instruction surfaces the raw arguments passed to bashTool, grepTool, and taskTool in the execution trace, allowing you to verify that the LLM generated correct parameters before the sandbox executes them.

Summary

  • Sandbox initialization: Wrap getSandbox calls in try-catch blocks to catch missing context errors early, as defined in packages/agent/tools/utils.ts.
  • Path resolution: Log the resolved workingDir in packages/agent/tools/bash.ts before calling sandbox.exec to catch directory traversal errors.
  • Result inspection: Check the success, exitCode, and truncated flags in sandbox results to distinguish between timeouts, errors, and output limits.
  • Approval flow: Override the needsApproval callback in bashTool to customize dangerous command detection and log rejection reasons.
  • Sub-agent tracing: Instrument UI handlers to log taskTool yields, monitoring pending states and startedAt timestamps to detect stuck agents.
  • Verbose logging: Inject debug instructions via customInstructions in packages/agent/system-prompt.ts to expose raw tool payloads.

Frequently Asked Questions

Why does my tool call fail with "Sandbox not initialized in context"?

This error originates in packages/agent/tools/utils.ts (lines 71-86) when the experimental_context passed to the tool execution does not contain a sandbox instance. Ensure your agent's prepareCall method initializes the context with experimental_context: { sandbox, ... } before invoking any tools.

How can I tell if a bash command timed out versus returned an error?

Inspect the sandbox result object returned by bashTool in packages/agent/tools/bash.ts. If success is false and stderr is empty, the command exceeded TIMEOUT_MS and was terminated. If success is false but stderr contains text, the command ran but exited with a non-zero status code.

What causes "Path traversal" errors in file operations?

The isPathWithinDirectory helper in packages/agent/tools/utils.ts (lines 23-33) validates that resolved paths stay within the sandbox root. This error occurs when the cwd parameter in bashTool or file paths in readTool/writeTool resolve to a location outside the allowed sandbox directory, often due to absolute paths or ../ sequences.

How do I debug a sub-agent that appears frozen or stuck in a loop?

Instrument your UI handler to log every yield from taskTool in packages/agent/tools/task.ts. If the stream shows continuous pending states without finish-step events, the sub-agent is likely waiting for an approval-required tool or stuck in a recursive loop. Check the startedAt timestamp to confirm how long the task has been running, and verify whether the pending tool name indicates a blocked bash command awaiting approval.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →