How to Implement DAG-Style Orchestration in Agent Systems
DAG-style orchestration in agent systems combines a static graph of task dependencies with an LLM-driven control loop that dynamically selects the next node to execute while maintaining resumable state through a unified event thread.
Directed Acyclic Graphs (DAGs) provide a robust framework for modeling complex agent workflows where each step depends on previous outputs without circular dependencies. According to the humanlayer/12-factor-agents repository, modern AI agents can leverage this structure while allowing the LLM to dynamically navigate the graph in real time rather than following hard-coded execution paths. This approach unifies the reliability of traditional workflow engines with the flexibility of autonomous decision-making.
Why DAG-Style Orchestration Matters for AI Agents
Traditional software orchestration treats the system as a static directed graph where each node executes in a predetermined order. As noted in content/brief-history-of-software.md, while "software is a directed graph," modern AI agents can selectively throw the DAG away—letting the LLM decide the execution path dynamically based on intermediate results rather than hard-coding every edge.
This hybrid approach offers two distinct advantages:
- Deterministic safety: Explicit dependency declarations prevent circular logic and ensure data flows correctly between transformations.
- Dynamic flexibility: The LLM can skip unnecessary nodes, reorder operations based on context, or pause for external input when the graph cannot proceed automatically.
Core Architecture of DAG-Style Agent Orchestration
Implementing this pattern requires three integrated layers that work together within the control loop.
Graph Definition
The graph definition layer establishes nodes (tasks) and edges (dependencies) as static metadata. Each node object contains a unique identifier, a callable implementation, and a list of downstream dependencies. This structure lives outside the LLM context and serves as guardrails for what can execute, not necessarily what will execute next.
Control-Flow Loop
The control-flow loop serves as the execution engine. In walkthrough/07-agent.py, the agent_loop function serializes the current execution state (the "thread") into JSON or XML and transmits it to the LLM via the DetermineNextStep prompt. The LLM returns an intent—such as run_node with a specific node ID—which the loop validates against the graph before execution.
This pattern appears in lines 38-55 of walkthrough/07-agent.py, where the loop repeatedly:
- Serializes the thread events.
- Requests the next action from the LLM.
- Executes the selected node if dependencies are satisfied.
- Appends the result as a new event to the thread.
State Persistence
The state persistence layer maintains a unified execution state (the Thread) that records every event—tool calls, results, human clarifications, and errors. As implemented in walkthrough/07-agent.py (lines 4-35), the Thread class appends immutable events to a list, creating a complete audit trail that allows the LLM to reason about the entire workflow history when deciding the next step.
Implementing the Agent Loop with Dependency Resolution
To build a concrete orchestrator, you combine the static graph with the dynamic loop. The following example demonstrates a DAG where data flows from fetch → analyze → notify, governed by the patterns found in walkthrough/07-agent.py.
# dag_orchestrator.py
import json
from typing import Callable, Dict, List, Optional
from dataclasses import dataclass, field
# ----------------------------------------------------------------------
# 1. Define the Graph Node Structure
# ----------------------------------------------------------------------
@dataclass
class Node:
nid: str
fn: Callable[[List[dict]], dict]
deps: List[str] = field(default_factory=list)
# ----------------------------------------------------------------------
# 2. Implement Task Functions
# ----------------------------------------------------------------------
def fetch_data(thread: List[dict]) -> dict:
"""Simulate data retrieval and append to thread."""
data = {"items": [1, 2, 3], "source": "api"}
return {"type": "tool_call", "data": {"tool": "fetch", "result": data}}
def analyze_data(thread: List[dict]) -> dict:
"""Analyze previously fetched data."""
fetch_event = next(
e for e in thread
if e["type"] == "tool_call" and e["data"]["tool"] == "fetch"
)
items = fetch_event["data"]["result"]["items"]
return {"type": "tool_call", "data": {"tool": "analyze", "result": {"sum": sum(items)}}}
def notify_user(thread: List[dict]) -> dict:
"""Send notification based on analysis."""
analysis = next(
e for e in thread
if e["type"] == "tool_call" and e["data"]["tool"] == "analyze"
)
return {"type": "notification", "data": {"message": f"Result is {analysis['data']['result']['sum']}"}}
# ----------------------------------------------------------------------
# 3. Static DAG Definition
# ----------------------------------------------------------------------
GRAPH = {
"fetch": Node("fetch", fetch_data),
"analyze": Node("analyze", analyze_data, deps=["fetch"]),
"notify": Node("notify", notify_user, deps=["analyze"]),
}
# ----------------------------------------------------------------------
# 4. Dependency Resolution
# ----------------------------------------------------------------------
def is_ready(node: Node, thread: List[dict]) -> bool:
"""Check if all dependencies appear as completed tool calls in thread."""
completed = {
e["data"]["tool"] for e in thread
if e["type"] == "tool_call"
}
return all(dep in completed for dep in node.deps)
# ----------------------------------------------------------------------
# 5. LLM-Driven Orchestration Loop
# ----------------------------------------------------------------------
def run_node(nid: str, thread: List[dict]) -> None:
"""Execute node and append result to thread."""
node = GRAPH[nid]
if not is_ready(node, thread):
raise ValueError(f"Dependencies not satisfied for {nid}")
result = node.fn(thread)
thread.append(result)
print(f"Executed {nid}: {result['data']}")
def determine_next_step(thread: List[dict]) -> Optional[str]:
"""
Simulates LLM decision logic.
In production, this calls the LLM with serialized thread.
"""
completed = {e["data"]["tool"] for e in thread if e["type"] == "tool_call"}
# Find first node with satisfied deps not yet completed
for nid, node in GRAPH.items():
if nid not in completed and is_ready(node, thread):
return nid
return None # DAG complete
def orchestrate():
"""Main execution loop mirroring walkthrough/07-agent.py patterns."""
thread = [{"type": "start", "data": "run_dag"}]
while True:
next_id = determine_next_step(thread)
if next_id is None:
print("DAG execution complete")
break
# In full implementation, LLM validates choice via agent_loop
run_node(next_id, thread)
if __name__ == "__main__":
orchestrate()
This implementation demonstrates the dynamic DAG concept: the static GRAPH defines legal transitions, while determine_next_step (standing in for the LLM via DetermineNextStep) decides which valid node to execute based on the current thread state.
Handling Asynchronous Interruptions and Human Approval
Real-world agent systems must handle interruptions—human approval requests, long-running async tasks, or error conditions. The content/factor-08-own-your-control-flow.md file describes how to break the execution loop safely without losing state.
When the LLM requires human input, the loop persists the partial thread and exits. Upon receiving the external input (via webhook or CLI), the system rehydrates the thread and resumes execution:
def orchestrate_with_interrupts(thread: List[dict]):
while True:
# Check if we're resuming from a human approval
if thread[-1].get("type") == "human_approval_request":
print("Waiting for external input...")
return # Exit loop, state is preserved in thread
next_id = determine_next_step(thread)
if next_id is None:
break
# Simulate LLM deciding it needs approval before notify
if next_id == "notify" and not any(e.get("approval") for e in thread):
thread.append({
"type": "human_approval_request",
"data": {"node": "notify", "reason": "Cost threshold exceeded"}
})
continue
run_node(next_id, thread)
This pattern, detailed in lines 23-45 and 55-68 of content/factor-08-own-your-control-flow.md, ensures that state lives in the thread, not in memory, making workflows resilient to crashes and restarts.
Summary
- DAG-style orchestration merges static dependency graphs with LLM-driven execution decisions, allowing agents to navigate complex workflows flexibly while maintaining data integrity.
- The Thread class in
walkthrough/07-agent.pyprovides the critical state persistence layer, serializing every event so the LLM can reason about the complete execution history when selecting the next node. - Dependency resolution occurs outside the LLM context— validating that the LLM's chosen node has satisfied prerequisites before execution.
- Interruptible control flow enables human-in-the-loop approvals and async operations by breaking the
agent_loopand resuming later without losing execution context, as documented incontent/factor-08-own-your-control-flow.md.
Frequently Asked Questions
How does DAG-style orchestration differ from traditional workflow engines?
Traditional workflow engines enforce rigid, pre-defined execution paths where each node triggers the next automatically. In DAG-style agent orchestration, the static graph defines possible transitions and dependencies, but the LLM decides which node to execute next based on the current thread state. This allows dynamic skipping, reordering, or conditional branching that rigid orchestrators cannot support without explicit condition nodes.
What happens when an agent workflow needs human approval?
When the LLM determines that human approval is required—either explicitly via tool design or implicitly through safety checks—the agent_loop breaks out of the control flow and persists the thread state. As implemented in content/factor-08-own-your-control-flow.md, the system stores the pending approval event in the thread and exits. Upon receiving the human response (via API webhook or CLI input), the process rehydrates the thread and resumes execution from the exact point of interruption.
Why is the Thread class essential for DAG orchestration?
The Thread class serves as the single source of truth for execution state. Unlike traditional workflow engines that track state implicitly in memory or database transactions, the thread explicitly records every tool call, result, and interruption as an immutable event. This unified execution state enables the LLM to make informed decisions about which DAG node to execute next, supports resumability across process restarts, and provides complete audit trails for debugging complex agent behaviors.
Can the LLM modify the DAG structure during execution?
While the LLM cannot modify the static graph definition (which acts as safety guardrails), it can effectively ignore branches by never selecting certain nodes via the DetermineNextStep intent. For true dynamic graph modification—adding nodes or edges at runtime—the system would need to extend the Node class and GRAPH structure, then validate changes against safety constraints before appending the new structure to the thread for subsequent LLM reasoning.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →