How the Egonex AI Graph-Reviewer Agent Validates Referential Integrity

The Egonex AI graph-reviewer agent validates referential integrity by generating a deterministic JavaScript script that verifies every edge, layer, and tour step references existing node IDs, outputting a JSON report that determines graph approval or rejection.

The Egonex AI graph-reviewer agent serves as a critical quality gate in the Understand-Anything repository, ensuring generated knowledge graphs contain no dangling references. When processing a KnowledgeGraph during Phase 6 of the /understand skill, the agent validates referential integrity through an automated script that performs exhaustive existence checks. This validation ensures that every edge connects real nodes and that layers and interactive tours only reference valid entities.

The Four-Step Validation Process

As defined in agents/graph-reviewer.md (lines 65-69), the agent implements Check 2 – Referential Integrity (Critical) via a script generated at runtime. The validation proceeds through four distinct phases to ensure complete graph consistency.

1. Node ID Index Construction

The script first builds a lookup map for O(1) existence verification by extracting all node identifiers from the nodes array:

const nodeIds = new Set(graph.nodes.map(n => n.id));

This Set structure enables constant-time lookups when validating references throughout the graph structure, providing optimal performance even for large knowledge graphs.

2. Edge Source and Target Validation

According to lines 65-66 of the agent specification, every edge undergoes dual verification to ensure both source and target properties reference existing nodes:

graph.edges.forEach((e, i) => {
  if (!nodeIds.has(e.source))
    issues.push(`Edge ${i} source '${e.source}' does not exist`);
  if (!nodeIds.has(e.target))
    issues.push(`Edge ${i} target '${e.target}' does not exist`);
});

Any edge referencing a non-existent node ID generates a precise error message containing the edge index and missing identifier.

3. Layer nodeIds Validation

Line 67 of the specification mandates checking each layer's nodeIds array to confirm all entries point to real nodes:

graph.layers?.forEach((layer, li) => {
  layer.nodeIds?.forEach(id => {
    if (!nodeIds.has(id))
      issues.push(`Layer ${li} contains missing nodeId '${id}'`);
  });
});

This prevents orphaned layer entries that would reference non-existent graph entities.

4. Tour Step nodeIds Validation

Line 68 requires that interactive tour steps undergo identical verification:

graph.tour?.steps?.forEach((step, si) => {
  step.nodeIds?.forEach(id => {
    if (!nodeIds.has(id))
      issues.push(`Tour step ${si} contains missing nodeId '${id}'`);
  });
});

Line 69 specifies that any dangling reference triggers issue reporting with precise location details.

Script Execution and Approval Logic

The agent writes the validation script to .understand-anything/tmp/ua-graph-validate.js and executes it via Node.js:

node .understand-anything/tmp/ua-graph-validate.js \
  "$PROJECT_ROOT/.understand-anything/knowledge-graph.json" \
  "$PROJECT_ROOT/.understand-anything/tmp/ua-review-results.json"

The script emits a JSON report (structured per lines 31-34) containing scriptCompleted: true, an issues array, and a warnings array. Per the graph-reviewer agent specification, the graph is rejected if the issues array contains any entries; only an empty array results in approval.

Example validation output showing referential integrity failures:

{
  "scriptCompleted": true,
  "issues": [
    "Edge 14 source 'file:src/missing.ts' does not exist",
    "Layer 2 contains missing nodeId 'function:util/unknown'"
  ],
  "warnings": []
}

Key Implementation Files

The referential integrity validation system spans several files in the Egonex-AI/Understand-Anything repository:

Summary

  • Deterministic validation: The Egonex AI graph-reviewer agent generates ua-graph-validate.js to perform exhaustive referential integrity checks with O(1) lookup performance
  • Four critical validation targets: Edge source/target properties (lines 65-66), layer nodeIds arrays (line 67), and tour step nodeIds arrays (line 68)
  • Zero-tolerance policy: Any dangling reference causes immediate graph rejection; only empty issues arrays result in approval
  • Machine-readable output: The validation produces a JSON report with scriptCompleted: true and detailed error messages consumable by downstream pipeline stages
  • Performance optimization: The script builds a Set of all node IDs before validation to ensure linear time complexity relative to the number of references

Frequently Asked Questions

What happens when the graph-reviewer detects a missing node reference?

The agent rejects the graph immediately. When the validation script finds a dangling reference in edges, layers, or tour steps, it adds a descriptive error to the issues array (line 69). Per the specification in agents/graph-reviewer.md, any non-empty issues array causes the agent to block approval, preventing corrupted knowledge graphs from advancing in the /understand skill pipeline.

How does the validation script achieve performance on large knowledge graphs?

The script constructs a Set containing all valid node IDs at initialization. This data structure provides O(1) lookup time for existence checks, enabling the script to validate thousands of references in linear time relative to the number of edges and layers, rather than quadratic time that would result from repeated array searches.

Can developers run the referential integrity check manually?

Yes. While the check runs automatically during Phase 6, the generated script at .understand-anything/tmp/ua-graph-validate.js is a standalone Node.js file. Developers can execute it manually with the knowledge graph JSON path and an output path as command-line arguments to debug referential integrity issues outside the agent workflow.

What distinguishes issues from warnings in the validation output?

The JSON report contains separate issues and warnings arrays (lines 31-34). Referential integrity violations always populate the issues array, which causes immediate graph rejection per the critical check designation. The warnings array accommodates non-critical concerns that don't block approval, though referential integrity failures are strictly treated as critical errors that halt the pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →