Sync SubAgent vs AsyncSubAgent: Performance and Concurrency Deep-Dive in DeepAgents
The synchronous SubAgentMiddleware blocks the main agent until the sub-agent finishes, while AsyncSubAgentMiddleware returns immediately with a job ID and processes work in the background on a remote LangGraph server.
DeepAgents provides two distinct middleware families for delegating tasks to sub-agents, each optimized for different latency and parallelism requirements. Understanding the performance and concurrency differences between sync SubAgent and AsyncSubAgent is critical for building responsive AI systems that handle both quick lookups and long-running background jobs. This analysis examines the implementation details found in the langchain-ai/deepagents repository to help you choose the right delegation pattern.
Execution Model: Blocking vs. Fire-and-Forget
The fundamental architectural difference lies in how each middleware handles the sub-agent lifecycle and return of control to the parent agent.
Synchronous SubAgent (Blocking)
In libs/deepagents/deepagents/middleware/subagents.py, the SubAgentMiddleware implements a blocking execution model. When the main agent invokes the task tool, it calls the sub-agent’s invoke method and waits synchronously for completion.
# Within the task tool (sync implementation)
def task(...):
subagent, subagent_state = _validate_and_prepare_state(...)
# Blocks until the sub‑agent finishes
result = subagent.invoke(subagent_state) # ← synchronous call
return _return_command_with_state_update(result, runtime.tool_call_id)
The _return_command_with_state_update logic (lines 445-447 in subagents.py) extracts the final message from the sub-agent’s messages list and performs a single atomic update to the parent state. The main thread remains blocked during the entire sub-agent execution, including token generation and any nested tool calls.
Asynchronous SubAgent (Non-Blocking)
The AsyncSubAgentMiddleware in libs/deepagents/deepagents/middleware/async_subagents.py implements a fire-and-forget pattern. Instead of waiting for completion, the middleware launches a LangGraph run on a remote server and returns a job ID instantly.
def launch_async_subagent(...):
# Immediately creates a LangGraph thread and run
client = clients.get_sync(subagent_type)
thread = client.threads.create()
run = client.runs.create(
thread_id=thread["thread_id"],
assistant_id=spec["graph_id"],
input={"messages": [{"role": "user", "content": description}]},
)
job_id = thread["thread_id"]
# Returns a job ID without waiting for the run to complete
return Command(
update={
"messages": [ToolMessage(f"Launched async subagent. job_id: {job_id}",
tool_call_id=runtime.tool_call_id)],
"async_subagent_jobs": {job_id: job},
}
)
This approach decouples job launch from result retrieval, allowing the main agent to continue processing other work while the sub-agent runs in isolation.
Latency and Responsiveness Comparison
Synchronous latency equals the total runtime of the sub-agent. If the sub-agent performs a 30-second research task, the main agent remains unresponsive for 30 seconds before receiving the result.
Asynchronous latency is limited strictly to the HTTP request overhead required to launch the job on the LangGraph server—typically milliseconds. The actual compute-heavy work happens in the background, so the main agent can deliver intermediate messages to the user immediately after launch.
This difference dramatically impacts perceived performance in user-facing applications where responsiveness is critical.
Concurrency Capabilities
Limited Parallelism in Sync SubAgent
The sync middleware supports parallel execution only through a specific pattern: issuing multiple task tool calls in a single LLM message. As noted in the middleware description (lines 37-38 in subagents.py), the LLM must emit several tool calls at once for any overlap to occur. Each call still blocks its own execution turn, so true concurrency is constrained by the LLM’s ability to generate parallel tool requests and the underlying runtime’s capacity to handle them simultaneously.
Unlimited Fire-and-Forget in Async SubAgent
The async middleware removes these constraints entirely. Because each launch_async_subagent call returns a job ID immediately without blocking, the main agent can fire off any number of async sub-agents in sequence. The documentation explicitly stresses this parallelism capability (lines 13-18 in async_subagents.py).
Jobs persist in the parent state under the async_subagent_jobs dictionary. You query status later using check_async_subagent or enumerate active jobs with list_async_subagent_jobs, enabling sophisticated workflow orchestration where dozens of background tasks proceed concurrently.
State Management and Error Handling
Synchronous state updates happen atomically upon completion. The final message from the sub-agent is extracted and merged into the parent state via _return_command_with_state_update, representing a single state transition.
Asynchronous state management uses a job tracking dictionary. Each Command object from the async middleware only modifies the async_subagent_jobs map and optionally appends a short acknowledgment message. Status checks return separate Command objects that update the parent agent with current progress.
Error handling diverges significantly:
- Sync: Errors raise inside the synchronous
invokecall and propagate immediately as part of the returnedToolMessage, halting the current tool execution. - Async: Errors are captured in the job’s status field (
error) and reported only when the user explicitly checks the job viacheck_async_subagent, allowing the main agent to handle failures asynchronously.
When to Use Each Approach
Choose SubAgentMiddleware (sync) for short-lived, deterministic tasks where the result is needed immediately to make the next decision. Ideal use cases include quick data lookups, small transformations, or validation steps that must complete before the parent agent can proceed.
Choose AsyncSubAgentMiddleware (async) for long-running, resource-heavy, or isolated tasks that should not block user interaction. This includes research tasks, code generation pipelines, data processing jobs, or any work that can proceed independently while the main agent handles other user queries or orchestrates additional parallel work.
Summary
- Sync SubAgent blocks the main thread until completion, limiting concurrency to parallel tool calls in a single turn and incurring full sub-agent runtime as latency.
- Async SubAgent returns job IDs immediately, enabling unlimited fire-and-forget concurrency with latency limited only to HTTP launch overhead.
- State handling differs atomically: sync updates the parent state once upon completion, while async maintains a persistent job tracking dictionary.
- Error propagation is immediate in sync (raising in
invoke) versus deferred in async (stored in job status). - Use sync for quick, blocking dependencies; use async for background processing and maximum throughput.
Frequently Asked Questions
How do I check the status of an async sub-agent?
Use the check_async_subagent tool with the job ID returned during launch. According to libs/deepagents/deepagents/middleware/async_subagents.py (lines 68-88), this function retrieves the run status from the remote LangGraph server using client.runs.get(), then builds a result object showing current progress or completion state.
Can I run multiple sync sub-agents in parallel?
Yes, but only if the LLM emits multiple task tool calls within a single message turn. The middleware description in subagents.py (lines 37-38) notes this pattern specifically. Each call still blocks individually, so true overlap requires the underlying execution environment to process the parallel tool calls simultaneously.
What happens if an async sub-agent fails?
Errors are captured in the job’s status field rather than raising immediately. When you call check_async_subagent, the middleware retrieves the run status and includes any error information in the returned result. This allows the main agent to decide whether to retry, cancel, or report the failure to the user without interrupting other concurrent operations.
Which middleware should I choose for long-running tasks?
Use AsyncSubAgentMiddleware. It launches jobs on remote LangGraph deployments with virtually no latency to the main agent, runs the work in the background, and supports status polling via check_async_subagent and list_async_subagent_jobs. This prevents the main agent from hanging during resource-intensive operations like research or code generation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →