How to Debug Open Agents Execution and Tool Call Issues: A Complete Guide
To debug Open Agents execution and tool call issues, inspect the sandbox initialization in utils.ts, validate working directory resolution in bash.ts, and trace the standardized execution result payload returned by the sandbox interface.
Open Agents by Vercel Labs executes every user request inside an isolated sandbox and routes actions through strictly typed tools. When a tool call fails, times out, or returns unexpected output, the debugging trace spans three distinct layers: the tool definition, the sandbox interface, and the runtime orchestration. This guide walks through the exact source code locations and instrumentation strategies needed to diagnose execution failures in the vercel-labs/open-agents repository.
Understanding the Open Agents Execution Architecture
Open Agents splits execution responsibility across three layers. Debugging requires checking each layer sequentially.
Tool Definition Layer
Typed tools reside in packages/agent/tools/*.ts and handle input validation, approval decisions, and sandbox invocation. For example, bashTool in packages/agent/tools/bash.ts executes shell commands, while grepTool and globTool handle file-system queries. Each tool imports getSandbox from utils.ts to access the isolated execution environment.
Sandbox Interface Layer
The sandbox interface is implemented by @open-harness/sandbox and exposed through helper functions in packages/agent/tools/utils.ts. All sandbox calls return a uniform result object containing success, exitCode, stdout, stderr, and an optional truncated flag. The getSandbox function (lines 71-86) validates that the sandbox exists in the experimental context and throws a descriptive error if initialization failed.
System Prompt and Runtime Orchestration
The core prompt logic in packages/agent/system-prompt.ts builds the agent context and injects model-specific overlays. This file also accepts customInstructions that can enable verbose logging, which surfaces debug information in the execution trace.
Debugging Sandbox Initialization Failures
The most common root cause of execution failures is a missing or uninitialized sandbox. The getSandbox function in packages/agent/tools/utils.ts throws an explicit error when the experimental context lacks a sandbox instance.
Wrap your tool execution in a try-catch block to surface the initialization error:
import { getSandbox } from "./utils";
export const bashTool = tool({
// ...
execute: async ({ command, cwd }, { experimental_context, abortSignal }) => {
let sandbox;
try {
sandbox = await getSandbox(experimental_context, "bash");
} catch (e) {
// Surface a clear error for the UI
return {
success: false,
exitCode: null,
stdout: "",
stderr: (e as Error).message,
};
}
// Proceed with execution...
const workingDir = cwd
? path.resolve(sandbox.workingDirectory, cwd)
: sandbox.workingDirectory;
const result = await sandbox.exec(command, workingDir, TIMEOUT_MS, {
signal: abortSignal,
});
return {
success: result.success,
exitCode: result.exitCode,
stdout: result.stdout,
stderr: result.stderr,
...(result.truncated && { truncated: true }),
};
},
});
If you see the error "Sandbox not initialized in context", verify that the agent's prepareCall function sets experimental_context: { sandbox, ... } before invoking the tool.
Resolving Working Directory and Path Issues
bashTool resolves the optional cwd parameter against the sandbox's workingDirectory (lines 8-14 in packages/agent/tools/bash.ts). A common bug is passing a relative path that resolves outside the sandbox root, which triggers a security error from isPathWithinDirectory (lines 23-33 in packages/agent/tools/utils.ts).
Log the resolved path before execution to catch traversal errors:
const workingDir = cwd
? path.isAbsolute(cwd) ? cwd : path.resolve(workingDirectory, cwd)
: workingDirectory;
console.log("[bashTool] resolved workingDir:", workingDir);
If the sandbox rejects the path, check that cwd is relative to the current task directory and not an absolute system path.
Inspecting Tool Execution Results
All sandbox executions return a standardized result object. When success is false and stderr is empty, the command likely timed out. When truncated is true, the output exceeded the ~50,000 character limit and was cut short.
Check the result shape in packages/agent/tools/bash.ts (lines 45-56):
return {
success: result.success,
exitCode: result.exitCode,
stdout: result.stdout,
stderr: result.stderr,
...(result.truncated && { truncated: true }),
};
Debug tips:
- Timeout: Increase
TIMEOUT_MSor optimize the command (e.g.,git log --oneline -n 20instead of full history). - Truncation: Pipe output through
heador use more specific search patterns to reduce verbosity.
Handling Dangerous Command Approvals
bashTool runs commandNeedsApproval (lines 31-47 in packages/agent/tools/bash.ts) to check against DANGEROUS_COMMAND_PATTERNS. If a pattern matches, the tool's needsApproval handler consults the optional options.needsApproval callback (lines 52-58).
When a command is rejected, the UI receives a needsApproval boolean. Investigate the caller that supplied options.needsApproval—usually a sub-agent or the parent agent implementing the approval policy.
Override the approval callback for granular control:
import { bashTool } from "packages/agent/tools/bash";
const safeBash = bashTool({
needsApproval: (args) => {
// Allow `rm -rf` only inside a temporary directory
if (/rm\s+-rf/.test(args.command)) {
return !args.cwd?.startsWith("tmp/");
}
return false; // other commands are safe
},
});
Debugging Sub-Agent Tasks
taskTool in packages/agent/tools/task.ts streams intermediate states back to the UI, including tool-call counts, pending calls, and token usage. The first yielded object contains a stable startedAt timestamp. Later yields contain pending with the sub-agent's internal tool name and raw input (line 21).
Instrument your UI handler to log each yield:
// Inside a UI component that receives taskTool streams
function handleTaskPart(part: TaskToolUIPart) {
if (part.pending) {
console.log(
"Sub-agent waiting on tool:",
part.pending.name,
"input:",
part.pending.input,
);
}
if (part.final) {
console.log("Sub-agent finished. Message:", part.final);
}
}
Debug tips:
- If yields show no
finish-stepevents, the sub-agent is stuck in a loop. - If
pendingnever becomesundefined, the sub-agent is waiting for a tool that requires approval.
Enabling Verbose Logging in the System Prompt
The core system prompt in packages/agent/system-prompt.ts accepts customInstructions (lines 28-33) that inject additional directives into the agent context. Add a debug instruction to force the agent to prepend diagnostic information to every tool payload.
Enable debug logging:
customInstructions: "Log every tool call payload and sandbox result."
This instruction surfaces the raw arguments passed to bashTool, grepTool, and taskTool in the execution trace, allowing you to verify that the LLM generated correct parameters before the sandbox executes them.
Summary
- Sandbox initialization: Wrap
getSandboxcalls in try-catch blocks to catch missing context errors early, as defined inpackages/agent/tools/utils.ts. - Path resolution: Log the resolved
workingDirinpackages/agent/tools/bash.tsbefore callingsandbox.execto catch directory traversal errors. - Result inspection: Check the
success,exitCode, andtruncatedflags in sandbox results to distinguish between timeouts, errors, and output limits. - Approval flow: Override the
needsApprovalcallback inbashToolto customize dangerous command detection and log rejection reasons. - Sub-agent tracing: Instrument UI handlers to log
taskToolyields, monitoringpendingstates andstartedAttimestamps to detect stuck agents. - Verbose logging: Inject debug instructions via
customInstructionsinpackages/agent/system-prompt.tsto expose raw tool payloads.
Frequently Asked Questions
Why does my tool call fail with "Sandbox not initialized in context"?
This error originates in packages/agent/tools/utils.ts (lines 71-86) when the experimental_context passed to the tool execution does not contain a sandbox instance. Ensure your agent's prepareCall method initializes the context with experimental_context: { sandbox, ... } before invoking any tools.
How can I tell if a bash command timed out versus returned an error?
Inspect the sandbox result object returned by bashTool in packages/agent/tools/bash.ts. If success is false and stderr is empty, the command exceeded TIMEOUT_MS and was terminated. If success is false but stderr contains text, the command ran but exited with a non-zero status code.
What causes "Path traversal" errors in file operations?
The isPathWithinDirectory helper in packages/agent/tools/utils.ts (lines 23-33) validates that resolved paths stay within the sandbox root. This error occurs when the cwd parameter in bashTool or file paths in readTool/writeTool resolve to a location outside the allowed sandbox directory, often due to absolute paths or ../ sequences.
How do I debug a sub-agent that appears frozen or stuck in a loop?
Instrument your UI handler to log every yield from taskTool in packages/agent/tools/task.ts. If the stream shows continuous pending states without finish-step events, the sub-agent is likely waiting for an approval-required tool or stuck in a recursive loop. Check the startedAt timestamp to confirm how long the task has been running, and verify whether the pending tool name indicates a blocked bash command awaiting approval.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →