How to Add Custom Middleware to the Deep Agent Execution Stack

You can extend the DeepAgents framework by subclassing AgentMiddleware, overriding the wrap_model_call or awrap_model_call hooks to transform ModelRequest objects, and passing the middleware instance to the middleware argument of create_deep_agent.

DeepAgents (langchain-ai/deepagents) provides a thin, extensible middleware layer that sits between the agent graph and the underlying LLM. This architecture allows you to intercept model requests, modify system prompts, filter available tools, or inject observability without altering core agent logic.

How the Middleware Stack Executes

The middleware system is assembled inside libs/deepagents/deepagents/graph.py within the create_deep_agent factory function (lines 38‑43). By default, DeepAgents applies a standard stack containing TodoListMiddleware, FilesystemMiddleware, SubAgentMiddleware, SummarizationMiddleware, AnthropicPromptCachingMiddleware, and PatchToolCallsMiddleware. Any custom middleware you provide is appended after these defaults, ensuring your transformations run last.

Each middleware class subclasses the generic AgentMiddleware type (re-exported from the LangChain agents package via deepagents.middleware.__init__) and implements one or both of these hooks:

Hook Signature Execution Path
wrap_model_call def wrap_model_call(self, request: ModelRequest[ContextT], handler: Callable[[ModelRequest[ContextT]], ModelResponse[ResponseT]]) -> ModelResponse[ResponseT] Synchronous agents
awrap_model_call async def awrap_model_call(self, request: ModelRequest[ContextT], handler: Callable[[ModelRequest[ContextT]], Awaitable[ModelResponse[ResponseT]]]) -> ModelResponse[ResponseT] Asynchronous agents

Request flow works as follows:

  1. The agent graph constructs a ModelRequest containing the system prompt, tools, and runtime state.
  2. The first middleware in the chain receives the request, optionally mutates it using request.override(), and calls handler(modified_request) to pass control downward.
  3. After the final middleware, the handler invokes the underlying LLM (model.invoke(request)).
  4. The ModelResponse bubbles back up through each middleware layer, allowing post-processing before returning to the agent.

Creating a Custom Middleware Class

To implement custom middleware, you must inherit from AgentMiddleware and override the synchronous hook, the asynchronous hook, or both.

from langchain.agents.middleware import AgentMiddleware
from langchain.agents.middleware.types import ModelRequest, ModelResponse
from typing import Callable, Awaitable

class LoggingMiddleware(AgentMiddleware):
    """Logs every model request and response for debugging."""

    def wrap_model_call(
        self,
        request: ModelRequest,
        handler: Callable[[ModelRequest], ModelResponse],
    ) -> ModelResponse:
        print(f"🟢 Outgoing request - Tools: {[t.get('name') for t in request.tools]}")
        
        # Forward to next middleware or LLM

        response = handler(request)
        
        print(f"🔵 Incoming response: {response}")
        return response

    async def awrap_model_call(
        self,
        request: ModelRequest,
        handler: Callable[[ModelRequest], Awaitable[ModelResponse]],
    ) -> ModelResponse:
        print(f"[async] Request tools: {len(request.tools)}")
        resp = await handler(request)
        print(f"[async] Response received")
        return resp

Critical implementation detail: The ModelRequest object is immutable. To modify it (e.g., prepending a system prompt), call request.override(**changes), which returns a new instance. For example, request.override(system_message=new_prompt).

Registering Middleware in create_deep_agent

Pass your middleware instance (or list) to the middleware kwarg when building a top-level agent:

from deepagents.graph import create_deep_agent
from example_middleware import LoggingMiddleware

agent = create_deep_agent(
    model="openai:gpt-4o-mini",
    middleware=[LoggingMiddleware()],  # Runs after default stack

)

Subagent-Specific Middleware

For subagents, include the middleware in the subagent specification's middleware field. The SubAgentMiddleware (lines 672‑682 in libs/deepagents/deepagents/middleware/subagents.py) merges user-provided middleware with the default subagent stack automatically:

sub_spec = {
    "name": "researcher",
    "description": "Performs web searches",
    "system_prompt": "You have access to the `search` tool.",
    "middleware": [LoggingMiddleware()],  # Only applies to this subagent

}

agent = create_deep_agent(
    model="anthropic:claude-3-5-sonnet",
    subagents=[sub_spec],
)

Real-World Patterns from the Source Code

Study the built-in middlewares for advanced patterns:

  • FilesystemMiddleware (libs/deepagents/deepagents/middleware/filesystem.py, lines 1100‑1145): Demonstrates dynamic system prompt construction by detecting backend capabilities and appending filesystem-specific instructions using append_to_system_message helpers.
  • SubAgentMiddleware (libs/deepagents/deepagents/middleware/subagents.py, lines 672‑682): Shows how to inspect the request context to determine if a subagent should be spawned and how to merge middleware stacks.
  • State Management: Middlewares like SummarizationMiddleware and SkillsMiddleware use PrivateStateAttr annotations (e.g., SkillsState.skills_metadata) to persist data across calls without polluting the public API.

Summary

  • DeepAgents exposes an AgentMiddleware base class (from LangChain agents) that lets you intercept LLM calls via wrap_model_call and awrap_model_call.
  • The stack is built in create_deep_agent (libs/deepagents/deepagents/graph.py), applying defaults first, then your custom middleware.
  • Use request.override() to transform immutable requests and handler(request) to forward control.
  • Register middleware at the top level via the middleware kwarg, or per-subagent via the subagent specification.

Frequently Asked Questions

Do I need to implement both sync and async hooks?

No. If you only plan to use synchronous agents, implementing wrap_model_call is sufficient. However, DeepAgents defaults to async execution in most production settings, so providing awrap_model_call ensures compatibility. You can usually have the async method call the sync implementation or vice versa if the logic is identical.

Can I remove or reorder the default middleware stack?

No, the default stack (TodoListMiddleware, FilesystemMiddleware, etc.) is automatically prepended inside create_deep_agent. Your custom middleware always executes after the defaults. If you need to disable a specific default behavior, you must subclass that middleware to create a no-op version and replace it in a forked version of the graph factory.

How do I persist state between middleware invocations?

Use private state attributes as demonstrated in SkillsMiddleware and SummarizationMiddleware. Annotate your state class with PrivateStateAttr, and store metadata there. The middleware receives the current state in the request object and can return modified state in the response, which the agent graph persists for the next turn.

Can middleware modify the LLM response before it reaches the agent?

Yes. While the primary pattern is request transformation, you can capture the response from handler(request), inspect or modify it, and return the modified object. This is useful for logging, redacting sensitive content, or implementing response caching layers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →