Implementing Subgraph Workflows and Reusable Workflow Modules in ChatDev

ChatDev implements subgraph workflows by treating reusable node collections as独立 directed graphs that can be embedded as single nodes within parent workflows, enabling modular, hierarchical AI agent pipelines.

In ChatDev, a workflow is fundamentally a directed graph where nodes represent agents, tools, or processes, and edges define execution flow. For complex projects requiring repetitive sub-processes—such as critique-revision cycles or validation steps—implementing subgraph workflows and reusable workflow modules in ChatDev prevents configuration duplication and maintains clean, maintainable YAML definitions. The subgraph mechanism allows developers to package entire graph structures into reusable components that execute within their own isolated contexts while remaining integrated with the parent workflow's state.

How Subgraphs Work in ChatDev

The subgraph architecture treats any node with type: subgraph as a container for an entire independent graph. When the GraphManager encounters such a node during graph construction, it delegates to GraphManager._build_subgraph() in workflow/graph_manager.py, which instantiates a nested GraphContext and recursively builds the internal graph structure.

Each subgraph operates within its own variable scope, memory stores, and log level configuration, defined in SubgraphInlineConfig.FIELD_SPECS within entity/configs/node/subgraph.py. This isolation ensures that internal subgraph operations do not pollute the parent graph's state unless explicitly passed through edge connections.

Core Components and Source Files

The implementation relies on four primary components that handle configuration, loading, and execution:

SubgraphConfig and Payload Classes — Located in entity/configs/node/subgraph.py (lines 22-126), these define the data models for subgraph declarations. SubgraphFileConfig handles external file references (lines 53-84), while SubgraphInlineConfig manages inline graph definitions (lines 75-88). The subgraph_source_registry (lines 22-45) maps type strings to their respective parser classes.

load_subgraph_config() — Found in workflow/subgraph_loader.py, this function reads external YAML files, resolves variable substitutions, and caches parsed graph dictionaries to prevent redundant disk I/O.

GraphManager._build_subgraph() — Implemented in workflow/graph_manager.py (lines 48-118), this method orchestrates subgraph construction by creating fresh GraphConfig instances and spawning recursive GraphManager instances for nested graph building.

GraphContext — Defined in workflow/graph_context.py (lines 15-33), this runtime container maintains the parent graph's state and manages a mapping of child subgraphs via graph.subgraphs[node_id], creating a tree structure of reusable modules.

Execution Flow

The subgraph execution process follows a hierarchical resolution pattern:

  1. Graph Parsing — GraphManager.build_graph_structure() iterates through the main YAML definition, identifying nodes where node_type equals "subgraph".

  2. Configuration Resolution — For each subgraph node, _build_subgraph(node_id) extracts the SubgraphConfig. If type is "config", it uses the inline dictionary directly; if "file", it invokes load_subgraph_config() to resolve the path relative to yaml_instance/, the parent directory, or the repository root.

  3. Variable Merging — The loaded graph dictionary merges with the parent's variable map (combined_vars), allowing subgraphs to inherit context while maintaining override capabilities.

  4. Recursive Construction — A new GraphManager instance builds the subgraph recursively, meaning subgraphs can contain their own subgraphs to arbitrary depth.

  5. Context Registration — The parent GraphContext stores the constructed child in graph.subgraphs[node_id], maintaining the hierarchical relationship.

  6. Execution Delegation — During runtime, GraphExecutor._execute_node() detects subgraph nodes, delegates execution to the child context, and returns resulting messages to the parent graph via standard edge connections.

Configuration Examples

Inline Subgraph Definition

For self-contained workflows where the subgraph logic is specific to the current graph, use the inline configuration method:


# yaml_instance/demo_sub_graph.yaml

graph:
  id: paper_gen
  start: [A]
  nodes:
    - id: A
      type: agent
      config:
        agent_name: Writer
        task: Draft paper
    - id: B
      type: subgraph          # Subgraph node declaration

      config:
        type: config          # Inline definition

        config:
          id: paper_critique
          nodes:
            - id: B1
              type: agent
              config:
                agent_name: Reviewer
                task: Check quality
          edges: []
    - id: C
      type: agent
      config:
        agent_name: Editor
  edges:
    - from: A
      to: B
    - from: B
      to: C

When GraphManager.build_graph() processes this definition, it automatically constructs the nested paper_critique graph and executes node B1 before proceeding to node C.

File-Based Reusable Modules

For cross-project reusability, externalize subgraph definitions to separate YAML files:


# parent_graph.yaml

graph:
  id: main_flow
  start: [Start]
  nodes:
    - id: Start
      type: agent
      config:
        agent_name: Initiator
    - id: Review
      type: subgraph
      config:
        type: file            # External file reference

        config:
          path: critique.yaml
    - id: Publish
      type: agent
      config:
        agent_name: Publisher
  edges:
    - from: Start
      to: Review
    - from: Review
      to: Publish

Place critique.yaml in yaml_instance/critique.yaml or specify an absolute path. This approach creates a single source of truth for the critique process, allowing multiple parent graphs to reference the same file without configuration duplication.

Python Implementation Guide

To programmatically build and execute graphs containing subgraphs:

from pathlib import Path
from workflow.graph_manager import GraphManager
from workflow.graph_context import GraphContext
from entity.graph_config import GraphConfig
from utils.io_utils import read_yaml

# Load the top-level design

design = read_yaml("yaml_instance/demo_sub_graph.yaml")
graph_cfg = GraphConfig.from_dict(
    config=design,
    name="demo_sub_graph",
    output_root=Path("./outputs"),
    source_path=str(Path("yaml_instance/demo_sub_graph.yaml").resolve()),
    vars={}
)

# Build runtime context and graph

graph_ctx = GraphContext(config=graph_cfg)
gm = GraphManager(graph_ctx)
gm.build_graph()                # Constructs subgraphs recursively

# Execute with a task prompt

from workflow.graph_executor import GraphExecutor
executor = GraphExecutor.execute_graph(
    graph=graph_ctx,
    task_prompt="Write a technical analysis of subgraph workflows."
)
print(executor.get_final_output())

The build_graph() call automatically resolves all subgraph nodes, loads external configurations, and constructs the complete execution tree before runtime begins.

Extending the Subgraph Registry

The subgraph_source_registry decouples source type identifiers from implementation classes, enabling custom subgraph sources without modifying core engine code. To add a database-backed subgraph loader:

from entity.configs.node.subgraph import register_subgraph_source
from mypkg.my_subgraph import MySubgraphConfig  # Subclass of BaseConfig

register_subgraph_source(
    name="mydb",
    config_cls=MySubgraphConfig,
    description="Load subgraph definition from a DB table"
)

Users can then reference the custom source in YAML:

type: mydb
config:
  query: "SELECT definition FROM subgraphs WHERE name='review_cycle'"

The engine's loading, variable resolution, and execution logic remain unchanged, as the registry abstraction handles instantiation.

Summary

  • Subgraph nodes in ChatDev encapsulate entire directed graphs within a single node declaration, enabling hierarchical workflow construction.
  • Two configuration modes exist: type: config for inline graph definitions and type: file for external YAML references that support cross-project reusability.
  • Isolation and inheritance are managed through GraphContext, where each subgraph maintains independent variables and memory while accessing parent context through merged variable maps.
  • Performance optimization occurs via load_subgraph_config()'s module-level _SUBGRAPH_CACHE, which prevents redundant file system operations when the same subgraph file is referenced multiple times.
  • Extensibility is achieved through subgraph_source_registry, allowing custom source types (databases, APIs, etc.) without altering the core execution engine in workflow/graph_manager.py or workflow/graph_executor.py.

Frequently Asked Questions

What is a subgraph workflow in ChatDev?

A subgraph workflow is a modular graph component defined in entity/configs/node/subgraph.py that allows a single node to contain an entire independent directed graph. When the GraphExecutor encounters a node with type: subgraph, it delegates execution to a nested GraphContext that runs the internal graph to completion before returning control to the parent workflow.

How do I reuse a workflow across multiple projects?

Define the reusable workflow in a separate YAML file within yaml_instance/ or an absolute path, then reference it using type: file in the subgraph configuration. The load_subgraph_config() function in workflow/subgraph_loader.py caches these definitions, ensuring that multiple parent graphs referencing the same file share a single source of truth without memory overhead.

Can subgraphs be nested?

Yes. The recursive implementation of GraphManager._build_subgraph() in workflow/graph_manager.py allows subgraphs to contain their own subgraph nodes, creating a tree structure of arbitrary depth. Each level maintains its own GraphContext instance stored in the parent's graph.subgraphs mapping.

How does subgraph caching work?

The workflow/subgraph_loader.py module maintains a module-level dictionary _SUBGRAPH_CACHE keyed by resolved file paths. When load_subgraph_config() is called, it checks this cache before reading the filesystem, ensuring that repeated references to the same subgraph file—whether from the same parent graph or different ones—do not trigger redundant I/O operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →