How to Add Custom Middleware to the Deep Agent Execution Stack
You can extend the DeepAgents framework by subclassing AgentMiddleware, overriding the wrap_model_call or awrap_model_call hooks to transform ModelRequest objects, and passing the middleware instance to the middleware argument of create_deep_agent.
DeepAgents (langchain-ai/deepagents) provides a thin, extensible middleware layer that sits between the agent graph and the underlying LLM. This architecture allows you to intercept model requests, modify system prompts, filter available tools, or inject observability without altering core agent logic.
How the Middleware Stack Executes
The middleware system is assembled inside libs/deepagents/deepagents/graph.py within the create_deep_agent factory function (lines 38‑43). By default, DeepAgents applies a standard stack containing TodoListMiddleware, FilesystemMiddleware, SubAgentMiddleware, SummarizationMiddleware, AnthropicPromptCachingMiddleware, and PatchToolCallsMiddleware. Any custom middleware you provide is appended after these defaults, ensuring your transformations run last.
Each middleware class subclasses the generic AgentMiddleware type (re-exported from the LangChain agents package via deepagents.middleware.__init__) and implements one or both of these hooks:
| Hook | Signature | Execution Path |
|---|---|---|
wrap_model_call |
def wrap_model_call(self, request: ModelRequest[ContextT], handler: Callable[[ModelRequest[ContextT]], ModelResponse[ResponseT]]) -> ModelResponse[ResponseT] |
Synchronous agents |
awrap_model_call |
async def awrap_model_call(self, request: ModelRequest[ContextT], handler: Callable[[ModelRequest[ContextT]], Awaitable[ModelResponse[ResponseT]]]) -> ModelResponse[ResponseT] |
Asynchronous agents |
Request flow works as follows:
- The agent graph constructs a
ModelRequestcontaining the system prompt, tools, and runtime state. - The first middleware in the chain receives the request, optionally mutates it using
request.override(), and callshandler(modified_request)to pass control downward. - After the final middleware, the handler invokes the underlying LLM (
model.invoke(request)). - The
ModelResponsebubbles back up through each middleware layer, allowing post-processing before returning to the agent.
Creating a Custom Middleware Class
To implement custom middleware, you must inherit from AgentMiddleware and override the synchronous hook, the asynchronous hook, or both.
from langchain.agents.middleware import AgentMiddleware
from langchain.agents.middleware.types import ModelRequest, ModelResponse
from typing import Callable, Awaitable
class LoggingMiddleware(AgentMiddleware):
"""Logs every model request and response for debugging."""
def wrap_model_call(
self,
request: ModelRequest,
handler: Callable[[ModelRequest], ModelResponse],
) -> ModelResponse:
print(f"🟢 Outgoing request - Tools: {[t.get('name') for t in request.tools]}")
# Forward to next middleware or LLM
response = handler(request)
print(f"🔵 Incoming response: {response}")
return response
async def awrap_model_call(
self,
request: ModelRequest,
handler: Callable[[ModelRequest], Awaitable[ModelResponse]],
) -> ModelResponse:
print(f"[async] Request tools: {len(request.tools)}")
resp = await handler(request)
print(f"[async] Response received")
return resp
Critical implementation detail: The ModelRequest object is immutable. To modify it (e.g., prepending a system prompt), call request.override(**changes), which returns a new instance. For example, request.override(system_message=new_prompt).
Registering Middleware in create_deep_agent
Pass your middleware instance (or list) to the middleware kwarg when building a top-level agent:
from deepagents.graph import create_deep_agent
from example_middleware import LoggingMiddleware
agent = create_deep_agent(
model="openai:gpt-4o-mini",
middleware=[LoggingMiddleware()], # Runs after default stack
)
Subagent-Specific Middleware
For subagents, include the middleware in the subagent specification's middleware field. The SubAgentMiddleware (lines 672‑682 in libs/deepagents/deepagents/middleware/subagents.py) merges user-provided middleware with the default subagent stack automatically:
sub_spec = {
"name": "researcher",
"description": "Performs web searches",
"system_prompt": "You have access to the `search` tool.",
"middleware": [LoggingMiddleware()], # Only applies to this subagent
}
agent = create_deep_agent(
model="anthropic:claude-3-5-sonnet",
subagents=[sub_spec],
)
Real-World Patterns from the Source Code
Study the built-in middlewares for advanced patterns:
FilesystemMiddleware(libs/deepagents/deepagents/middleware/filesystem.py, lines 1100‑1145): Demonstrates dynamic system prompt construction by detecting backend capabilities and appending filesystem-specific instructions usingappend_to_system_messagehelpers.SubAgentMiddleware(libs/deepagents/deepagents/middleware/subagents.py, lines 672‑682): Shows how to inspect the request context to determine if a subagent should be spawned and how to merge middleware stacks.- State Management: Middlewares like
SummarizationMiddlewareandSkillsMiddlewareusePrivateStateAttrannotations (e.g.,SkillsState.skills_metadata) to persist data across calls without polluting the public API.
Summary
- DeepAgents exposes an
AgentMiddlewarebase class (from LangChain agents) that lets you intercept LLM calls viawrap_model_callandawrap_model_call. - The stack is built in
create_deep_agent(libs/deepagents/deepagents/graph.py), applying defaults first, then your custom middleware. - Use
request.override()to transform immutable requests andhandler(request)to forward control. - Register middleware at the top level via the
middlewarekwarg, or per-subagent via the subagent specification.
Frequently Asked Questions
Do I need to implement both sync and async hooks?
No. If you only plan to use synchronous agents, implementing wrap_model_call is sufficient. However, DeepAgents defaults to async execution in most production settings, so providing awrap_model_call ensures compatibility. You can usually have the async method call the sync implementation or vice versa if the logic is identical.
Can I remove or reorder the default middleware stack?
No, the default stack (TodoListMiddleware, FilesystemMiddleware, etc.) is automatically prepended inside create_deep_agent. Your custom middleware always executes after the defaults. If you need to disable a specific default behavior, you must subclass that middleware to create a no-op version and replace it in a forked version of the graph factory.
How do I persist state between middleware invocations?
Use private state attributes as demonstrated in SkillsMiddleware and SummarizationMiddleware. Annotate your state class with PrivateStateAttr, and store metadata there. The middleware receives the current state in the request object and can return modified state in the response, which the agent graph persists for the next turn.
Can middleware modify the LLM response before it reaches the agent?
Yes. While the primary pattern is request transformation, you can capture the response from handler(request), inspect or modify it, and return the modified object. This is useful for logging, redacting sensitive content, or implementing response caching layers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →