How DeepAgents Implements Tool Retries and Error Recovery During Agent Execution

DeepAgents implements a tiered "detect-classify-recover" strategy using middleware layers including summarization fallback, timeout guidance, large-result eviction, and synthetic error injection to handle tool failures without breaking agent control flow.

The langchain-ai/deepagents repository provides a robust execution framework for AI agents that implements comprehensive tool retries and error recovery mechanisms. Rather than terminating sessions on failures, the architecture wraps tool execution in middleware layers that detect exceptions, classify error types, and execute context-specific recovery strategies to maintain agent continuity.

Automatic Recovery from Context Overflow

When model calls exceed token budgets, the SummarizationMiddleware in libs/deepagents/deepagents/middleware/summarization.py (lines 880-984) implements an immediate fallback strategy. Instead of raising a hard failure, the middleware detects the ContextOverflowError, classifies it as a token budget issue, and recovers by off-loading old messages, generating a concise summary, and retrying the model call with the summary plus recent messages.

from deepagents.deepagents.middleware.summarization import SummarizationMiddleware

# Assume `request` is a ModelRequest that would overflow the token budget.

try:
    response = model.invoke(request)          # may raise ContextOverflowError

except ContextOverflowError:
    # Summarisation middleware automatically:

    #  - off‑loads old messages,

    #  - creates a summary,

    #  - retries the model call with the summary + recent messages.

    response = summarizer.wrap_model_call(request, model.invoke)

This mechanism ensures agents can continue long-running conversations without manual intervention when context windows fill up.

Timeout Error Guidance and Retry Prompting

The LocalShellBackend handles command timeouts by returning structured error messages that guide the LLM toward recovery. As implemented in the test suite at libs/deepagents/tests/unit_tests/test_local_shell.py (lines 81-88), when a subprocess times out, the backend suggests using the timeout parameter on subsequent calls, effectively prompting the agent to retry with adjusted parameters rather than failing permanently.

backend = LocalShellBackend(timeout=1, inherit_env=True)

with patch("subprocess.run", side_effect=subprocess.TimeoutExpired("cmd", 1)):
    result = backend.execute("sleep 10")
    # Output contains: “… timed out … use the timeout parameter …”

    print(result.output)

This approach transforms transient infrastructure limits into actionable guidance for the language model.

Graceful Handling of Oversized Tool Results

For tool outputs exceeding tool_token_limit_before_evict, the FilesystemMiddleware in libs/deepagents/deepagents/middleware/filesystem.py (lines 124-165) implements large-result eviction to prevent context overflow. The middleware persists oversized results to backend storage, replaces inline content with a preview, and continues execution. If the eviction write fails, the system logs the error but still returns a usable preview, allowing the agent to retry the operation later rather than crashing.


# Inside FilesystemMiddleware._create_read_file_tool()

if token_limit and len(content) >= NUM_CHARS_PER_TOKEN * token_limit:
    # Evict to backend, replace with preview + file reference

    file_path = backend.write_file(...)
    return TOO_LARGE_TOOL_MSG.format(tool_call_id=runtime.tool_call_id,
                                    file_path=file_path,
                                    content_sample=_create_content_preview(content))

This eviction strategy ensures that reading large files never breaks the agent's control flow.

Synthetic Error Injection for Dangling Tool Calls

When tool calls are cancelled or dropped by external systems, PatchToolCallsMiddleware in libs/deepagents/deepagents/middleware/patch_tool_calls.py (lines 24-43) ensures deterministic failure handling. The middleware scans message histories for AIMessage instances containing tool calls that lack corresponding ToolMessage responses, then injects synthetic error messages so the agent perceives the failure and can choose to retry or take alternative actions.


# In PatchToolCallsMiddleware.before_agent

if isinstance(msg, AIMessage) and msg.tool_calls:
    for tool_call in msg.tool_calls:
        if not any(
            t for t in messages[i:]
            if t.type == "tool" and t.tool_call_id == tool_call["id"]
        ):
            # Insert synthetic error ToolMessage so the agent can retry

            patched_messages.append(
                ToolMessage(
                    content=f"Tool call {tool_call['name']} with id {tool_call['id']} was cancelled …",
                    name=tool_call["name"],
                    tool_call_id=tool_call["id"],
                )
            )

This patching mechanism prevents agents from hanging indefinitely while waiting for tool results that will never arrive.

Structured Error Handling for Async Sub-Agents

The async sub-agent middleware in libs/deepagents/deepagents/middleware/async_subagents.py (lines 316-322 and 590-597) prevents silent failure swallowing in distributed execution scenarios. When remote LangGraph servers fail or return unknown statuses, the middleware returns structured error objects containing an error field with clear messages, enabling the main agent to implement appropriate retry logic or fallback behaviors rather than proceeding with incomplete data.

Prevention of Infinite Retry Loops

DeepAgents implements explicit safeguards against infinite retry loops through cutoff tracking and cache invalidation strategies.

The LocalContextMiddleware in libs/cli/deepagents_cli/local_context.py (lines 18-22) records the _local_context_refreshed_at_cutoff index after summarization events to prevent redundant detection script execution. If the script fails after a context refresh, the system records the cutoff and stops further retries of the same failed operation.

if raw_event is not None:
    cutoff = event.get("cutoff_index")
    refreshed = state.get("_local_context_refreshed_at_cutoff")
    if cutoff != refreshed:
        output = self._run_detect_script()
        if output:
            return {"local_context": output,
                    "_local_context_refreshed_at_cutoff": cutoff}
        # Record the cutoff to stop further retries

        return {"_local_context_refreshed_at_cutoff": cutoff}

Additionally, configuration loaders in libs/cli/deepagents_cli/model_config.py (line 1327) explicitly avoid caching error responses, ensuring that subsequent invocations attempt fresh operations—useful when transient issues like permission errors are resolved externally.

Finally, the execution graph in libs/deepagents/deepagents/graph.py (line 63) includes advisory comments recommending workflow abortion when steps fail repeatedly, encouraging diagnostic analysis over blind retry attempts.

Summary

  • Context overflow triggers automatic summarization and retry via SummarizationMiddleware (lines 880-984).
  • Timeout errors return actionable guidance prompting parameter adjustment in subsequent calls.
  • Oversized results are evicted to storage with previews returned, preventing context breakage (lines 124-165).
  • Dangling tool calls are patched with synthetic error messages to unblock agent decision-making (lines 24-43).
  • Async failures are surfaced as structured errors rather than silent omissions (lines 316-322, 590-597).
  • Infinite loops are prevented via cutoff tracking and explicit non-caching of error states.

Frequently Asked Questions

How does DeepAgents handle token limit exceeded errors?

When a model call triggers a ContextOverflowError, the SummarizationMiddleware in libs/deepagents/deepagents/middleware/summarization.py (lines 880-984) intercepts the exception, compresses historical messages into a summary, and retries the call with the condensed context. This allows agents to continue execution without manual context management.

What happens when a shell command times out?

The LocalShellBackend captures TimeoutExpired exceptions and returns error messages containing specific guidance to use the timeout parameter on subsequent attempts. As shown in libs/deepagents/tests/unit_tests/test_local_shell.py (lines 81-88), this effectively prompts the LLM to retry with longer timeouts rather than failing permanently.

How does DeepAgents prevent infinite retry loops?

The system employs multiple safeguards: the LocalContextMiddleware tracks cutoff indices via _local_context_refreshed_at_cutoff to prevent redundant script execution (lines 18-22), configuration loaders avoid caching error responses (line 1327), and the execution graph recommends aborting workflows that fail repeatedly (line 63).

What mechanism handles tool calls that never return results?

The PatchToolCallsMiddleware in libs/deepagents/deepagents/middleware/patch_tool_calls.py (lines 24-43) scans for AIMessage tool calls lacking corresponding ToolMessage responses. It injects synthetic error messages into the conversation history, allowing the agent to perceive the failure and decide whether to retry or pursue alternative actions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →