# How the Egonex AI Graph-Reviewer Agent Validates Referential Integrity

> Discover how the Egonex AI graph-reviewer agent validates referential integrity. Generates a script to verify all graph elements then outputs a JSON report for approval or rejection.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-20

---

**The Egonex AI graph-reviewer agent validates referential integrity by generating a deterministic JavaScript script that verifies every edge, layer, and tour step references existing node IDs, outputting a JSON report that determines graph approval or rejection.**

The **Egonex AI** **graph-reviewer** agent serves as a critical quality gate in the **Understand-Anything** repository, ensuring generated knowledge graphs contain no dangling references. When processing a `KnowledgeGraph` during Phase 6 of the `/understand` skill, the agent validates referential integrity through an automated script that performs exhaustive existence checks. This validation ensures that every edge connects real nodes and that layers and interactive tours only reference valid entities.

## The Four-Step Validation Process

As defined in [`agents/graph-reviewer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/agents/graph-reviewer.md) (lines 65-69), the agent implements **Check 2 – Referential Integrity (Critical)** via a script generated at runtime. The validation proceeds through four distinct phases to ensure complete graph consistency.

### 1. Node ID Index Construction

The script first builds a lookup map for O(1) existence verification by extracting all node identifiers from the `nodes` array:

```javascript
const nodeIds = new Set(graph.nodes.map(n => n.id));

```

This Set structure enables constant-time lookups when validating references throughout the graph structure, providing optimal performance even for large knowledge graphs.

### 2. Edge Source and Target Validation

According to lines 65-66 of the agent specification, every edge undergoes dual verification to ensure both `source` and `target` properties reference existing nodes:

```javascript
graph.edges.forEach((e, i) => {
  if (!nodeIds.has(e.source))
    issues.push(`Edge ${i} source '${e.source}' does not exist`);
  if (!nodeIds.has(e.target))
    issues.push(`Edge ${i} target '${e.target}' does not exist`);
});

```

Any edge referencing a non-existent node ID generates a precise error message containing the edge index and missing identifier.

### 3. Layer nodeIds Validation

Line 67 of the specification mandates checking each layer's `nodeIds` array to confirm all entries point to real nodes:

```javascript
graph.layers?.forEach((layer, li) => {
  layer.nodeIds?.forEach(id => {
    if (!nodeIds.has(id))
      issues.push(`Layer ${li} contains missing nodeId '${id}'`);
  });
});

```

This prevents orphaned layer entries that would reference non-existent graph entities.

### 4. Tour Step nodeIds Validation

Line 68 requires that interactive tour steps undergo identical verification:

```javascript
graph.tour?.steps?.forEach((step, si) => {
  step.nodeIds?.forEach(id => {
    if (!nodeIds.has(id))
      issues.push(`Tour step ${si} contains missing nodeId '${id}'`);
  });
});

```

Line 69 specifies that any dangling reference triggers issue reporting with precise location details.

## Script Execution and Approval Logic

The agent writes the validation script to [`.understand-anything/tmp/ua-graph-validate.js`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/tmp/ua-graph-validate.js) and executes it via Node.js:

```bash
node .understand-anything/tmp/ua-graph-validate.js \
  "$PROJECT_ROOT/.understand-anything/knowledge-graph.json" \
  "$PROJECT_ROOT/.understand-anything/tmp/ua-review-results.json"

```

The script emits a JSON report (structured per lines 31-34) containing `scriptCompleted: true`, an `issues` array, and a `warnings` array. Per the `graph-reviewer` agent specification, the graph is **rejected** if the `issues` array contains any entries; only an empty array results in **approval**.

Example validation output showing referential integrity failures:

```json
{
  "scriptCompleted": true,
  "issues": [
    "Edge 14 source 'file:src/missing.ts' does not exist",
    "Layer 2 contains missing nodeId 'function:util/unknown'"
  ],
  "warnings": []
}

```

## Key Implementation Files

The referential integrity validation system spans several files in the **Egonex-AI/Understand-Anything** repository:

- **[`agents/graph-reviewer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/agents/graph-reviewer.md)** – Defines Check 2 (Referential Integrity) at lines 65-69 and specifies that non-empty `issues` arrays trigger rejection
- **[`skills/understand/graph-reviewer-prompt.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/skills/understand/graph-reviewer-prompt.md)** – Provides the prompt template used when the `--review` flag forces LLM-based analysis
- **[`docs/superpowers/specs/2026-04-09-understand-knowledge-design.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/docs/superpowers/specs/2026-04-09-understand-knowledge-design.md)** – Design specification detailing the agent's integration into the Phase 6 pipeline

## Summary

- **Deterministic validation**: The Egonex AI graph-reviewer agent generates [`ua-graph-validate.js`](https://github.com/Egonex-AI/Understand-Anything/blob/main/ua-graph-validate.js) to perform exhaustive referential integrity checks with O(1) lookup performance
- **Four critical validation targets**: Edge source/target properties (lines 65-66), layer `nodeIds` arrays (line 67), and tour step `nodeIds` arrays (line 68)
- **Zero-tolerance policy**: Any dangling reference causes immediate graph rejection; only empty `issues` arrays result in approval
- **Machine-readable output**: The validation produces a JSON report with `scriptCompleted: true` and detailed error messages consumable by downstream pipeline stages
- **Performance optimization**: The script builds a Set of all node IDs before validation to ensure linear time complexity relative to the number of references

## Frequently Asked Questions

### What happens when the graph-reviewer detects a missing node reference?

The agent rejects the graph immediately. When the validation script finds a dangling reference in edges, layers, or tour steps, it adds a descriptive error to the `issues` array (line 69). Per the specification in [`agents/graph-reviewer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/agents/graph-reviewer.md), any non-empty `issues` array causes the agent to block approval, preventing corrupted knowledge graphs from advancing in the `/understand` skill pipeline.

### How does the validation script achieve performance on large knowledge graphs?

The script constructs a `Set` containing all valid node IDs at initialization. This data structure provides O(1) lookup time for existence checks, enabling the script to validate thousands of references in linear time relative to the number of edges and layers, rather than quadratic time that would result from repeated array searches.

### Can developers run the referential integrity check manually?

Yes. While the check runs automatically during Phase 6, the generated script at [`.understand-anything/tmp/ua-graph-validate.js`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/tmp/ua-graph-validate.js) is a standalone Node.js file. Developers can execute it manually with the knowledge graph JSON path and an output path as command-line arguments to debug referential integrity issues outside the agent workflow.

### What distinguishes issues from warnings in the validation output?

The JSON report contains separate `issues` and `warnings` arrays (lines 31-34). Referential integrity violations always populate the `issues` array, which causes immediate graph rejection per the critical check designation. The `warnings` array accommodates non-critical concerns that don't block approval, though referential integrity failures are strictly treated as critical errors that halt the pipeline.