How to Merge Subdomain Knowledge Graphs (Frontend/Backend) in Understand-Anything

Use the merge-subdomain-graphs.py utility script to automatically discover, deduplicate, and combine multiple subdomain knowledge graphs into a single unified knowledge-graph.json file.

When analyzing large codebases with distinct architectural boundaries, Lum1104/Understand-Anything generates separate knowledge graphs for each subdomain (such as frontend and backend). The framework includes a deterministic merge process that unifies these subgraphs while resolving entity conflicts, enabling the dashboard to visualize the entire system architecture as one cohesive graph.

Where Subdomain Knowledge Graphs Are Stored

Understand-Anything persists knowledge graph data in the .understand-anything/ directory at your project root. When scanning subdomains independently, the system produces separate JSON files following the naming pattern *knowledge-graph*.json:

The unified output is written to .understand-anything/knowledge-graph.json, which serves as the single source of truth for the dashboard and downstream skills. According to the core persistence layer in understand-anything-plugin/packages/core/src/persistence/index.ts, the constant GRAPH_FILE = "knowledge-graph.json" defines this target filename.

How the Merge Process Works

The merge functionality resides in understand-anything-plugin/skills/understand/merge-subdomain-graphs.py. This script implements a six-stage pipeline that guarantees deterministic output regardless of file discovery order.

File Discovery

The script walks $PROJECT_ROOT/.understand-anything/ and collects every JSON file matching the glob *knowledge-graph*.json while explicitly excluding knowledge-graph.json itself. This auto-discovery eliminates manual file enumeration.

Base Graph Loading

If a pre-existing knowledge-graph.json is present, it loads first to serve as the base graph. This preserves shared utilities and global types that span multiple subdomains.

Deduplication and Conflict Resolution

The script uses Map objects keyed by entity IDs to eliminate duplicates. Nodes and edges are processed with a subdomain-wins strategy:

  1. Subdomain graphs are inserted at the beginning of the processing array
  2. The base graph is appended last
  3. Later entries overwrite earlier ones in the Map

This ensures that specific subdomain definitions override generic base definitions, preserving the most granular view of the code. Edge deduplication uses a composite key of (source, target, type) to identify unique relationships.

Running the Merge Manually

While the merge step runs automatically as part of the understand skill workflow, you can execute it manually for custom pipelines or debugging:


# From your project root

python understand-anything-plugin/skills/understand/merge-subdomain-graphs.py $PROJECT_ROOT

Omitting the file list argument triggers auto-discovery of all subdomain graphs in the .understand-anything/ directory.

Python Implementation Details

The merge logic relies on dictionary-based deduplication to achieve O(n) complexity. Here is the core implementation pattern from merge-subdomain-graphs.py:

import json
import pathlib
import sys

project_root = pathlib.Path(sys.argv[1])
ua_dir = project_root / ".understand-anything"

# 1. Discover subdomain files

graph_files = [
    p for p in ua_dir.iterdir()
    if p.suffix == ".json"
    and p.name != "knowledge-graph.json"
    and "knowledge-graph" in p.name
]

# 2. Load existing base graph (if present)

graphs = []
base_path = ua_dir / "knowledge-graph.json"
if base_path.is_file():
    graphs.append(json.load(base_path.open()))

# 3. Prepend subdomain graphs so they take precedence

for p in graph_files:
    graphs.insert(0, json.load(p.open()))

# 4. Merge with deduplication

node_map = {}
edge_map = {}
for g in graphs:
    for n in g.get("nodes", []):
        node_map[n["id"]] = n  # Later assignments overwrite earlier ones

    for e in g.get("edges", []):
        edge_map[(e["source"], e["target"], e.get("type"))] = e

merged = {
    "nodes": list(node_map.values()),
    "edges": list(edge_map.values()),
    "kind": "knowledge",
}

# 5. Write unified graph

output_path = ua_dir / "knowledge-graph.json"
output_path.write_text(json.dumps(merged, indent=2))
print(f"Merged {len(graph_files)} graphs → {len(merged['nodes'])} nodes, {len(merged['edges'])} edges")

The script validates output against the Zod schema defined in understand-anything-plugin/packages/core/src/types.ts to ensure compatibility with the dashboard consumption layer.

Consuming the Unified Graph

After merging, the dashboard automatically fetches the unified graph from the development server. In understand-anything-plugin/packages/dashboard/src/App.tsx, the data URL constructor targets the merged file:

const dataUrl = (file: string, token?: string) =>
  `${import.meta.env.VITE_GRAPH_URL}/${file}${token ? `?token=${encodeURIComponent(token)} : ''}`;

fetch(dataUrl("knowledge-graph.json", accessToken))
  .then(r => r.json())
  .then(setGraph);

This ensures that visualization includes both frontend and backend entities with their cross-domain relationships intact.

Summary

  • Subdomain graphs are stored as separate JSON files matching *knowledge-graph*.json in the .understand-anything/ directory.
  • merge-subdomain-graphs.py automates discovery and unification, implementing a deterministic subdomain-wins conflict resolution strategy.
  • Deduplication uses Maps keyed by node ID and edge composite keys, with subdomain data inserted before base data to ensure specificity prevails.
  • Manual execution is supported by passing the project root as a command-line argument to the Python script.
  • Dashboard integration consumes the unified knowledge-graph.json automatically, providing a complete architectural view across all subdomains.

Frequently Asked Questions

What naming convention identifies subdomain knowledge graphs?

The merge script discovers any file in .understand-anything/ containing the substring knowledge-graph with a .json extension, excluding only the final knowledge-graph.json output file. This allows flexible naming like frontend-knowledge-graph.json or api-knowledge-graph.json.

How does the merge script handle conflicting node definitions?

When the same node ID appears in multiple subgraphs, the version from the subdomain graph takes precedence. The script achieves this by inserting subdomain graphs at the beginning of the processing array and the base graph at the end, ensuring that Map assignments from the base graph are overwritten by subdomain entries.

Can I run the merge outside the automated skill workflow?

Yes. Invoke the script directly with python merge-subdomain-graphs.py $PROJECT_ROOT. This is useful for CI/CD pipelines or when manually curating subgraphs from external sources before visualization.

Where does the dashboard load the unified graph from?

The dashboard fetches /knowledge-graph.json relative to the VITE_GRAPH_URL environment variable, as implemented in packages/dashboard/src/App.tsx. This endpoint serves the merged file produced by the consolidation script.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →