# How to Add Custom Middleware to the Deep Agent Execution Stack

> Learn to add custom middleware to the Deep Agent execution stack by subclassing AgentMiddleware and overriding hooks. Enhance your agent's functionality today.

- Repository: [LangChain/deepagents](https://github.com/langchain-ai/deepagents)
- Tags: how-to-guide
- Published: 2026-03-17

---

**You can extend the DeepAgents framework by subclassing `AgentMiddleware`, overriding the `wrap_model_call` or `awrap_model_call` hooks to transform `ModelRequest` objects, and passing the middleware instance to the `middleware` argument of `create_deep_agent`.**

DeepAgents (langchain-ai/deepagents) provides a thin, extensible middleware layer that sits between the agent graph and the underlying LLM. This architecture allows you to intercept model requests, modify system prompts, filter available tools, or inject observability without altering core agent logic.

## How the Middleware Stack Executes

The middleware system is assembled inside [`libs/deepagents/deepagents/graph.py`](https://github.com/langchain-ai/deepagents/blob/main/libs/deepagents/deepagents/graph.py) within the `create_deep_agent` factory function (lines 38‑43). By default, DeepAgents applies a standard stack containing `TodoListMiddleware`, `FilesystemMiddleware`, `SubAgentMiddleware`, `SummarizationMiddleware`, `AnthropicPromptCachingMiddleware`, and `PatchToolCallsMiddleware`. Any custom middleware you provide is appended **after** these defaults, ensuring your transformations run last.

Each middleware class subclasses the generic `AgentMiddleware` type (re-exported from the LangChain agents package via `deepagents.middleware.__init__`) and implements one or both of these hooks:

| Hook | Signature | Execution Path |
|------|-----------|--------------|
| `wrap_model_call` | `def wrap_model_call(self, request: ModelRequest[ContextT], handler: Callable[[ModelRequest[ContextT]], ModelResponse[ResponseT]]) -> ModelResponse[ResponseT]` | Synchronous agents |
| `awrap_model_call` | `async def awrap_model_call(self, request: ModelRequest[ContextT], handler: Callable[[ModelRequest[ContextT]], Awaitable[ModelResponse[ResponseT]]]) -> ModelResponse[ResponseT]` | Asynchronous agents |

**Request flow works as follows:**

1. The agent graph constructs a `ModelRequest` containing the system prompt, tools, and runtime state.
2. The first middleware in the chain receives the request, optionally mutates it using `request.override()`, and calls `handler(modified_request)` to pass control downward.
3. After the final middleware, the handler invokes the underlying LLM (`model.invoke(request)`).
4. The `ModelResponse` bubbles back up through each middleware layer, allowing post-processing before returning to the agent.

## Creating a Custom Middleware Class

To implement custom middleware, you must inherit from `AgentMiddleware` and override the synchronous hook, the asynchronous hook, or both.

```python
from langchain.agents.middleware import AgentMiddleware
from langchain.agents.middleware.types import ModelRequest, ModelResponse
from typing import Callable, Awaitable

class LoggingMiddleware(AgentMiddleware):
    """Logs every model request and response for debugging."""

    def wrap_model_call(
        self,
        request: ModelRequest,
        handler: Callable[[ModelRequest], ModelResponse],
    ) -> ModelResponse:
        print(f"🟢 Outgoing request - Tools: {[t.get('name') for t in request.tools]}")
        
        # Forward to next middleware or LLM

        response = handler(request)
        
        print(f"🔵 Incoming response: {response}")
        return response

    async def awrap_model_call(
        self,
        request: ModelRequest,
        handler: Callable[[ModelRequest], Awaitable[ModelResponse]],
    ) -> ModelResponse:
        print(f"[async] Request tools: {len(request.tools)}")
        resp = await handler(request)
        print(f"[async] Response received")
        return resp

```

**Critical implementation detail:** The `ModelRequest` object is immutable. To modify it (e.g., prepending a system prompt), call `request.override(**changes)`, which returns a new instance. For example, `request.override(system_message=new_prompt)`.

## Registering Middleware in create_deep_agent

Pass your middleware instance (or list) to the `middleware` kwarg when building a top-level agent:

```python
from deepagents.graph import create_deep_agent
from example_middleware import LoggingMiddleware

agent = create_deep_agent(
    model="openai:gpt-4o-mini",
    middleware=[LoggingMiddleware()],  # Runs after default stack

)

```

### Subagent-Specific Middleware

For subagents, include the middleware in the subagent specification's `middleware` field. The `SubAgentMiddleware` (lines 672‑682 in [`libs/deepagents/deepagents/middleware/subagents.py`](https://github.com/langchain-ai/deepagents/blob/main/libs/deepagents/deepagents/middleware/subagents.py)) merges user-provided middleware with the default subagent stack automatically:

```python
sub_spec = {
    "name": "researcher",
    "description": "Performs web searches",
    "system_prompt": "You have access to the `search` tool.",
    "middleware": [LoggingMiddleware()],  # Only applies to this subagent

}

agent = create_deep_agent(
    model="anthropic:claude-3-5-sonnet",
    subagents=[sub_spec],
)

```

## Real-World Patterns from the Source Code

Study the built-in middlewares for advanced patterns:

- **`FilesystemMiddleware`** ([`libs/deepagents/deepagents/middleware/filesystem.py`](https://github.com/langchain-ai/deepagents/blob/main/libs/deepagents/deepagents/middleware/filesystem.py), lines 1100‑1145): Demonstrates dynamic system prompt construction by detecting backend capabilities and appending filesystem-specific instructions using `append_to_system_message` helpers.
- **`SubAgentMiddleware`** ([`libs/deepagents/deepagents/middleware/subagents.py`](https://github.com/langchain-ai/deepagents/blob/main/libs/deepagents/deepagents/middleware/subagents.py), lines 672‑682): Shows how to inspect the request context to determine if a subagent should be spawned and how to merge middleware stacks.
- **State Management**: Middlewares like `SummarizationMiddleware` and `SkillsMiddleware` use `PrivateStateAttr` annotations (e.g., `SkillsState.skills_metadata`) to persist data across calls without polluting the public API.

## Summary

- **DeepAgents** exposes an `AgentMiddleware` base class (from LangChain agents) that lets you intercept LLM calls via `wrap_model_call` and `awrap_model_call`.
- The stack is built in `create_deep_agent` ([`libs/deepagents/deepagents/graph.py`](https://github.com/langchain-ai/deepagents/blob/main/libs/deepagents/deepagents/graph.py)), applying defaults first, then your custom middleware.
- Use `request.override()` to transform immutable requests and `handler(request)` to forward control.
- Register middleware at the top level via the `middleware` kwarg, or per-subagent via the subagent specification.

## Frequently Asked Questions

### Do I need to implement both sync and async hooks?

No. If you only plan to use synchronous agents, implementing `wrap_model_call` is sufficient. However, DeepAgents defaults to async execution in most production settings, so providing `awrap_model_call` ensures compatibility. You can usually have the async method call the sync implementation or vice versa if the logic is identical.

### Can I remove or reorder the default middleware stack?

No, the default stack (`TodoListMiddleware`, `FilesystemMiddleware`, etc.) is automatically prepended inside `create_deep_agent`. Your custom middleware always executes **after** the defaults. If you need to disable a specific default behavior, you must subclass that middleware to create a no-op version and replace it in a forked version of the graph factory.

### How do I persist state between middleware invocations?

Use **private state attributes** as demonstrated in `SkillsMiddleware` and `SummarizationMiddleware`. Annotate your state class with `PrivateStateAttr`, and store metadata there. The middleware receives the current state in the `request` object and can return modified state in the response, which the agent graph persists for the next turn.

### Can middleware modify the LLM response before it reaches the agent?

Yes. While the primary pattern is request transformation, you can capture the response from `handler(request)`, inspect or modify it, and return the modified object. This is useful for logging, redacting sensitive content, or implementing response caching layers.