# Graph-Reviewer Agent Validation: Ensuring Knowledge Graph Completeness and Integrity

> Discover how the graph-reviewer agent ensures knowledge graph completeness and integrity via a two-phase validation process enforcing schema compliance, referential integrity, and structural completeness.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: tutorial
- Published: 2026-06-12

---

**The graph-reviewer agent validates Knowledge Graphs through a deterministic two-phase process that enforces schema compliance, referential integrity, and structural completeness, rejecting any graph containing critical issues while surfacing warnings for quality improvements.**

The graph-reviewer agent in the Egonex-AI/Understand-Anything repository serves as the quality gate for Knowledge Graphs generated by the analysis pipeline. According to the source specification in [`agents/graph-reviewer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/agents/graph-reviewer.md), the agent executes a Node.js validation script that performs concrete checks for completeness and integrity before a final decision step approves or rejects the graph.

## Two-Phase Validation Architecture

The validation workflow is split into two distinct phases. Phase 1 runs a deterministic Node.js script that performs all concrete checks, while Phase 2 makes the final approval decision based on the script's output.

### Phase 1: Deterministic Script Execution

During Phase 1, the agent generates and executes [`.understand-anything/tmp/ua-graph-validate.js`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/tmp/ua-graph-validate.js). This script parses the Knowledge Graph JSON and executes a series of validation functions, populating a report object with any critical issues or warnings encountered. The script is coordinated by [`agents/assemble-reviewer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/agents/assemble-reviewer.md), which handles execution and retries if the script crashes.

### Phase 2: Approval Decision

Phase 2 reads the intermediate report from [`.understand-anything/intermediate/review.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/review.json). The agent approves the graph only when the `issues` array is empty. Any critical issue forces an immediate rejection, while warnings are preserved for diagnostic purposes. The final validation JSON is written back to the same path without the internal `scriptCompleted` flag.

## Critical Integrity Checks

The script enforces several critical integrity constraints that prevent structural failures in downstream consumption.

### Schema Validation

Every node must contain the fields `id`, `type`, `name`, `summary`, `tags`, and `complexity` with correct types. Every edge must contain `source`, `target`, `type`, `direction`, and `weight`. Missing or incorrectly typed fields generate critical issues that prevent approval.

### Referential Integrity

The validator ensures all edge `source` and `target` IDs reference existing nodes. Additionally, all `nodeIds` referenced in layers and tour steps must point to valid nodes. Dangling pointers create critical issues that break the connected graph structure.

### Uniqueness Constraints

Duplicate node IDs are not allowed. The script maintains a `Set` of seen IDs and reports any duplicates as critical issues, preventing ambiguous references that would break both the validator and the UI.

### Layer Coverage Requirements

File-level nodes—including types `file`, `config`, `document`, `service`, `pipeline`, `table`, `schema`, `resource`, and `endpoint`—must appear in exactly one layer's `nodeIds` array. Layers with empty `nodeIds` arrays also trigger critical issues, ensuring the visual dashboard fully represents all file-level entities.

## Completeness Validation Rules

Beyond integrity, the validator confirms the graph represents a substantive analysis rather than an empty or partial result.

### Minimum Structural Requirements

The graph must contain at least one node, one edge, one layer, and one tour step. If any of these structural elements are missing, the script appends a critical issue to the report.

### Domain Node Exceptions

When the graph contains domain nodes (`domain`, `flow`, or `step`), the requirements for layers and tour steps are relaxed to warnings only. This accommodates high-level conceptual graphs that may not require full layer decomposition or guided tours.

## Quality and Warning Checks

Non-critical checks generate warnings rather than rejections, encouraging higher-quality graphs without blocking the pipeline.

### Tour Step Validation

Tour steps must have sequential `order` values starting at 1 with no duplicates. Each step must reference at least one node, and the total step count must fall between 5 and 15. Violations are flagged as warnings.

### Content Quality Standards

Summaries must be non-empty and not merely the filename. The script also warns on self-referencing edges and reports orphan nodes (nodes with zero incident edges) to improve human readability.

### Non-Code Node Requirements

Specific node types such as `document`, `service`, `pipeline`, `table`, `schema`, `domain`, and `flow` are expected to have particular edge types. Missing expected edges generates warnings to encourage richer semantic connections.

### Naming Convention Compliance

The node's `type` must match the prefix of its `id` (for example, `type: "config"` requires an `id` starting with `config:`). This check ensures predictable naming conventions that simplify downstream tooling.

## Validation Report Structure

The script always outputs a JSON report with three top-level fields consumed by Phase 2:

- `issues` – An array of critical problems including schema violations, broken references, and duplicate IDs.
- `warnings` – An array of non-critical observations such as orphan nodes or missing expected edges.
- `stats` – A summary object containing node counts, edge counts, type distributions, and tour step totals.

## Implementation in the Validation Script

The validation logic is implemented in [`.understand-anything/tmp/ua-graph-validate.js`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/tmp/ua-graph-validate.js). Below is the skeleton showing key validation functions:

```javascript
const fs = require('fs');
const path = process.argv[2];
const out = process.argv[3];

const graph = JSON.parse(fs.readFileSync(path, 'utf8'));
const report = {
  scriptCompleted: true,
  issues: [],
  warnings: [],
  stats: {}
};

// Schema validation
function checkNodeSchema(node, idx) {
  const required = ['id', 'type', 'name', 'summary', 'tags', 'complexity'];
  required.forEach(f => {
    if (!(f in node)) report.issues.push(`Node ${idx} missing field "${f}"`);
  });
}
graph.nodes.forEach(checkNodeSchema);

// Referential integrity
function checkEdgeRefs(edge, idx) {
  const src = graph.nodes.find(n => n.id === edge.source);
  const tgt = graph.nodes.find(n => n.id === edge.target);
  if (!src) report.issues.push(`Edge ${idx} source "${edge.source}" does not exist`);
  if (!tgt) report.issues.push(`Edge ${idx} target "${edge.target}" does not exist`);
}
graph.edges.forEach(checkEdgeRefs);

// Completeness checks
if (graph.nodes.length === 0) report.issues.push('Graph contains no nodes');
if (graph.edges.length === 0) report.issues.push('Graph contains no edges');
if (graph.layers?.length === 0) report.issues.push('Graph contains no layers');

// Layer coverage
graph.layers?.forEach((layer, li) => {
  if (!layer.nodeIds?.length) report.issues.push(`Layer ${li} has empty nodeIds`);
});

// Uniqueness validation
const seen = new Set();
graph.nodes.forEach((n, i) => {
  if (seen.has(n.id)) report.issues.push(`Duplicate node id "${n.id}" at index ${i}`);
  else seen.add(n.id);
});

// Statistics compilation
report.stats = {
  totalNodes: graph.nodes.length,
  totalEdges: graph.edges.length,
  totalLayers: (graph.layers || []).length,
  tourSteps: (graph.tour?.steps || []).length,
  nodeTypes: graph.nodes.reduce((acc, {type}) => {
    acc[type] = (acc[type] || 0) + 1;
    return acc;
  }, {}),
  edgeTypes: graph.edges.reduce((acc, {type}) => {
    acc[type] = (acc[type] || 0) + 1;
    return acc;
  }, {})
};

fs.writeFileSync(out, JSON.stringify(report, null, 2));
process.exit(0);

```

Phase 2 then processes this report:

```javascript
const result = JSON.parse(fs.readFileSync('.understand-anything/intermediate/review.json'));
const final = {
  approved: result.issues.length === 0,
  issues: result.issues,
  warnings: result.warnings,
  stats: result.stats
};
fs.writeFileSync('.understand-anything/intermediate/review.json', JSON.stringify(final));

```

## Summary

- The graph-reviewer agent performs a two-phase validation: Phase 1 executes a deterministic Node.js script, and Phase 2 approves or rejects based on the results.
- Critical integrity checks include schema validation, referential integrity for all edge and layer references, uniqueness constraints for node IDs, and layer coverage requirements for file-level nodes.
- Completeness validation ensures the graph contains at least one node, edge, layer, and tour step, with relaxed rules when domain nodes are present.
- Quality checks generate warnings for tour sequencing issues, orphan nodes, content quality problems, and missing semantic edges on non-code nodes.
- The validation script outputs a structured JSON report with `issues`, `warnings`, and `stats` fields, stored at [`.understand-anything/intermediate/review.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/review.json).

## Frequently Asked Questions

### What constitutes a critical issue versus a warning in graph-reviewer validation?

Critical issues include schema violations, broken referential integrity (dangling edge references or missing layer nodeIds), duplicate node IDs, and missing required structural elements. These prevent graph approval. Warnings cover quality improvements such as orphan nodes, non-sequential tour steps, missing expected edges on non-code nodes, and deviations from naming conventions.

### Where does the graph-reviewer agent store its validation results?

The agent writes the intermediate validation report to [`.understand-anything/intermediate/review.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/intermediate/review.json). This file contains the `issues`, `warnings`, and `stats` arrays, along with the final `approved` boolean flag indicating whether the graph passed Phase 2 validation.

### How does the validation script handle domain nodes differently?

When the graph contains domain-level nodes such as `domain`, `flow`, or `step`, the requirements for layers and tour steps are downgraded from critical issues to warnings. This allows high-level conceptual graphs to pass validation even without full layer decomposition or guided tours.

### What specific fields are required for nodes and edges in the Knowledge Graph?

According to the specification in [`agents/graph-reviewer.md`](https://github.com/Egonex-AI/Understand-Anything/blob/main/agents/graph-reviewer.md), every node must include `id`, `type`, `name`, `summary`, `tags`, and `complexity`. Every edge must include `source`, `target`, `type`, `direction`, and `weight`. Missing any of these fields or providing incorrect types generates a critical validation issue.