# How the Graph-Reviewer Validates Completeness and Referential Integrity of Knowledge Graphs in Understand-Anything

> Discover how the graph-reviewer validates knowledge graph completeness and referential integrity in Understand-Anything. Learn about its deterministic pre-validation process.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-06

---

**The graph-reviewer agent performs deterministic pre-validation by executing a generated Node.js script that verifies every edge, layer, and tour step references existing nodes, while enforcing minimum structural requirements (nodes, edges, layers, tour steps) before any LLM-driven review begins.**

The `graph-reviewer` in the Lum1104/Understand-Anything repository is a deterministic quality assurance step that validates knowledge graph integrity as a prerequisite to LLM analysis. According to the agent specification in [`understand-anything-plugin/agents/graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/graph-reviewer.md), this component generates and executes hard-coded validation scripts to enforce strict referential integrity and structural completeness across the graph’s nodes, edges, and metadata layers.

## Referential Integrity Checks

The validator ensures that every ID referenced within the graph actually exists in the `nodes` array. These checks prevent dangling pointers that would break downstream analysis.

### Edge Source and Target Validation

Every relationship in the graph must connect two valid entities. The script iterates through the `edges` array and validates that both endpoints exist:

- **Edge source exists**: Every `source` field must reference a node ID present in the `nodes` array.
- **Edge target exists**: Every `target` field must reference a node ID present in the `nodes` array.

If either check fails, the script logs a critical issue specifying the exact edge index and the missing ID (e.g., *"Edge at index 14 references non-existent target node 'file:src/missing.ts'"*), as defined in lines 63-66 and 69-70 of [`graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-reviewer.md).

### Layer and Tour Step Node References

Beyond edges, the validator checks secondary references in structural metadata:

- **Layer node references**: Each entry in `nodeIds` inside a layer object must point to a valid node ID.
- **Tour step node references**: Each `nodeIds` entry within tour steps must reference an existing node.

These validations ensure that visualization layers and guided tours do not point to removed or non-existent entities (lines 66-68).

## Completeness Validation

The graph-reviewer enforces minimum structural requirements to ensure the graph contains sufficient data for meaningful analysis.

### Minimum Structural Requirements

Unless the graph is classified as a domain graph, the following are mandatory:

- **At least one node**: The `nodes` array must contain ≥ 1 entry.
- **At least one edge**: The `edges` array must contain ≥ 1 entry.
- **At least one layer**: Structural graphs must define ≥ 1 layer.
- **At least one tour step**: Structural graphs must define ≥ 1 tour step.

Missing any of these critical components results in an entry in the `issues` array, forcing the final decision to *Rejected* (lines 71-77).

### Domain Graph Detection Logic

The validator relaxes layer and tour step requirements when analyzing domain-specific graphs. If any node has a type of `domain`, `flow`, or `step`, the script treats the graph as a domain graph and downgrades missing layers or tour steps from critical errors to warnings (lines 78-79). This distinction allows high-level conceptual graphs to pass validation without requiring detailed structural tours.

## Validation Execution and Reporting

The graph-reviewer operates in two deterministic phases that require no LLM inference unless explicitly requested via the `--review` flag.

### Phase 1: Script Generation and Execution

The agent writes a validation script (Node.js by default) to `$PROJECT_ROOT/.understand-anything/tmp/ua-graph-validate.js`. This script reads the graph JSON and executes the referential and completeness checks described above.

```bash

# Execute the deterministic validation script

node $PROJECT_ROOT/.understand-anything/tmp/ua-graph-validate.js \
  "<path-to-graph-json>" \
  "$PROJECT_ROOT/.understand-anything/tmp/ua-review-results.json"

```

### Phase 2: Result Parsing and Decision Logic

After execution, the agent parses the generated JSON output file containing `issues`, `warnings`, and `stats`:

```json
{
  "scriptCompleted": true,
  "issues": [
    "Edge at index 14 references non-existent target node 'file:src/missing.ts'"
  ],
  "warnings": [
    "3 function nodes have no edges connecting to them"
  ],
  "stats": { "totalNodes": 42, "totalEdges": 87 }
}

```

The final decision logic is binary: if the `issues` array is empty, the agent sets `approved: true`; otherwise, the graph is rejected. All critical integrity failures discovered in Phase 1 block progression to any optional LLM review phase.

## Summary

- The graph-reviewer is a **deterministic pre-check** that runs before LLM analysis in the Understand-Anything pipeline.
- **Referential integrity** is enforced by verifying that every `source`, `target`, `layer.nodeIds`, and `tourStep.nodeIds` references existing nodes.
- **Completeness** requires at least one node, one edge, one layer, and one tour step for structural graphs.
- **Domain graphs** (containing `domain`, `flow`, or `step` nodes) receive relaxed requirements where missing layers or tours generate warnings instead of critical errors.
- The validation script outputs a JSON report with `issues` (critical) and `warnings` (non-critical) arrays; any issue results in rejection.
- Source specifications are located in [`understand-anything-plugin/agents/graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/graph-reviewer.md).

## Frequently Asked Questions

### What happens if an edge references a non-existent node?

The validation script logs a critical issue detailing the specific edge index and the missing node ID, adds this to the `issues` array in the output JSON, and marks the graph as rejected. This prevents downstream components from processing broken relationships.

### How does the graph-reviewer handle domain-specific graphs differently?

When the validator detects any node with type `domain`, `flow`, or `step`, it classifies the graph as a domain graph. In this mode, missing layers or tour steps are downgraded from critical errors to warnings, allowing high-level conceptual maps to pass validation without exhaustive structural metadata.

### Where is the validation logic defined in the Understand-Anything repository?

The hard-coded validation rules are specified in [`understand-anything-plugin/agents/graph-reviewer.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/agents/graph-reviewer.md) (lines 63-79), which defines the deterministic checks for referential integrity and completeness. The optional LLM-powered review prompt is stored separately in [`understand-anything-plugin/skills/understand/graph-reviewer-prompt.md`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/skills/understand/graph-reviewer-prompt.md) and is only used when the `--review` flag is passed.

### Can the deterministic validation be bypassed to use only LLM review?

No. The deterministic validation is a mandatory pre-step in the standard pipeline. The LLM review phase (`--review` flag) is optional and only executes *after* the deterministic script completes successfully. Any critical issues found by the deterministic validator will reject the graph before it reaches any LLM analysis.