# How to Merge Subdomain Knowledge Graphs (Frontend/Backend) in Understand-Anything

> Easily merge frontend and backend subdomain knowledge graphs into one unified graph using the merge-subdomain-graphs.py utility in Understand Anything. Automate discovery and deduplication.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-05-31

---

**Use the [`merge-subdomain-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-subdomain-graphs.py) utility script to automatically discover, deduplicate, and combine multiple subdomain knowledge graphs into a single unified [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) file.**

When analyzing large codebases with distinct architectural boundaries, Lum1104/Understand-Anything generates separate knowledge graphs for each subdomain (such as frontend and backend). The framework includes a deterministic merge process that unifies these subgraphs while resolving entity conflicts, enabling the dashboard to visualize the entire system architecture as one cohesive graph.

## Where Subdomain Knowledge Graphs Are Stored

Understand-Anything persists knowledge graph data in the `.understand-anything/` directory at your project root. When scanning subdomains independently, the system produces separate JSON files following the naming pattern `*knowledge-graph*.json`:

- [`frontend-knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/frontend-knowledge-graph.json)
- [`backend-knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/backend-knowledge-graph.json)

The unified output is written to [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json), which serves as the single source of truth for the dashboard and downstream skills. According to the core persistence layer in [`understand-anything-plugin/packages/core/src/persistence/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/persistence/index.ts), the constant `GRAPH_FILE = "knowledge-graph.json"` defines this target filename.

## How the Merge Process Works

The merge functionality resides in **[`understand-anything-plugin/skills/understand/merge-subdomain-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/merge-subdomain-graphs.py)**. This script implements a six-stage pipeline that guarantees deterministic output regardless of file discovery order.

### File Discovery

The script walks `$PROJECT_ROOT/.understand-anything/` and collects every JSON file matching the glob `*knowledge-graph*.json` while explicitly excluding [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) itself. This auto-discovery eliminates manual file enumeration.

### Base Graph Loading

If a pre-existing [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) is present, it loads first to serve as the **base graph**. This preserves shared utilities and global types that span multiple subdomains.

### Deduplication and Conflict Resolution

The script uses `Map` objects keyed by entity IDs to eliminate duplicates. Nodes and edges are processed with a **subdomain-wins** strategy:

1. Subdomain graphs are inserted at the beginning of the processing array
2. The base graph is appended last
3. Later entries overwrite earlier ones in the Map

This ensures that specific subdomain definitions override generic base definitions, preserving the most granular view of the code. Edge deduplication uses a composite key of `(source, target, type)` to identify unique relationships.

## Running the Merge Manually

While the merge step runs automatically as part of the `understand` skill workflow, you can execute it manually for custom pipelines or debugging:

```bash

# From your project root

python understand-anything-plugin/skills/understand/merge-subdomain-graphs.py $PROJECT_ROOT

```

Omitting the file list argument triggers auto-discovery of all subdomain graphs in the `.understand-anything/` directory.

## Python Implementation Details

The merge logic relies on dictionary-based deduplication to achieve O(n) complexity. Here is the core implementation pattern from [`merge-subdomain-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-subdomain-graphs.py):

```python
import json
import pathlib
import sys

project_root = pathlib.Path(sys.argv[1])
ua_dir = project_root / ".understand-anything"

# 1. Discover subdomain files

graph_files = [
    p for p in ua_dir.iterdir()
    if p.suffix == ".json"
    and p.name != "knowledge-graph.json"
    and "knowledge-graph" in p.name
]

# 2. Load existing base graph (if present)

graphs = []
base_path = ua_dir / "knowledge-graph.json"
if base_path.is_file():
    graphs.append(json.load(base_path.open()))

# 3. Prepend subdomain graphs so they take precedence

for p in graph_files:
    graphs.insert(0, json.load(p.open()))

# 4. Merge with deduplication

node_map = {}
edge_map = {}
for g in graphs:
    for n in g.get("nodes", []):
        node_map[n["id"]] = n  # Later assignments overwrite earlier ones

    for e in g.get("edges", []):
        edge_map[(e["source"], e["target"], e.get("type"))] = e

merged = {
    "nodes": list(node_map.values()),
    "edges": list(edge_map.values()),
    "kind": "knowledge",
}

# 5. Write unified graph

output_path = ua_dir / "knowledge-graph.json"
output_path.write_text(json.dumps(merged, indent=2))
print(f"Merged {len(graph_files)} graphs → {len(merged['nodes'])} nodes, {len(merged['edges'])} edges")

```

The script validates output against the Zod schema defined in [`understand-anything-plugin/packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/types.ts) to ensure compatibility with the dashboard consumption layer.

## Consuming the Unified Graph

After merging, the dashboard automatically fetches the unified graph from the development server. In [`understand-anything-plugin/packages/dashboard/src/App.tsx`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/dashboard/src/App.tsx), the data URL constructor targets the merged file:

```typescript
const dataUrl = (file: string, token?: string) =>
  `${import.meta.env.VITE_GRAPH_URL}/${file}${token ? `?token=${encodeURIComponent(token)} : ''}`;

fetch(dataUrl("knowledge-graph.json", accessToken))
  .then(r => r.json())
  .then(setGraph);

```

This ensures that visualization includes both frontend and backend entities with their cross-domain relationships intact.

## Summary

- **Subdomain graphs** are stored as separate JSON files matching `*knowledge-graph*.json` in the `.understand-anything/` directory.
- **[`merge-subdomain-graphs.py`](https://github.com/Lum1104/Understand-Anything/blob/main/merge-subdomain-graphs.py)** automates discovery and unification, implementing a deterministic subdomain-wins conflict resolution strategy.
- **Deduplication** uses Maps keyed by node ID and edge composite keys, with subdomain data inserted before base data to ensure specificity prevails.
- **Manual execution** is supported by passing the project root as a command-line argument to the Python script.
- **Dashboard integration** consumes the unified [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) automatically, providing a complete architectural view across all subdomains.

## Frequently Asked Questions

### What naming convention identifies subdomain knowledge graphs?

The merge script discovers any file in `.understand-anything/` containing the substring `knowledge-graph` with a `.json` extension, excluding only the final [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json) output file. This allows flexible naming like [`frontend-knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/frontend-knowledge-graph.json) or [`api-knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/api-knowledge-graph.json).

### How does the merge script handle conflicting node definitions?

When the same node ID appears in multiple subgraphs, the version from the subdomain graph takes precedence. The script achieves this by inserting subdomain graphs at the beginning of the processing array and the base graph at the end, ensuring that Map assignments from the base graph are overwritten by subdomain entries.

### Can I run the merge outside the automated skill workflow?

Yes. Invoke the script directly with `python merge-subdomain-graphs.py $PROJECT_ROOT`. This is useful for CI/CD pipelines or when manually curating subgraphs from external sources before visualization.

### Where does the dashboard load the unified graph from?

The dashboard fetches [`/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main//knowledge-graph.json) relative to the `VITE_GRAPH_URL` environment variable, as implemented in [`packages/dashboard/src/App.tsx`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/dashboard/src/App.tsx). This endpoint serves the merged file produced by the consolidation script.