# How Graph Schema Validation Ensures Referential Integrity in Egonex-AI/Understand-Anything

> Learn how Egonex-AI/Understand-Anything uses graph schema validation via the validateGraph function to enforce referential integrity, ensuring all node references are valid and preventing data inconsistencies.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: internals
- Published: 2026-06-14

---

**The `validateGraph` function in [`schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/schema.ts) enforces referential integrity by verifying that every edge's `source` and `target` properties reference existing node IDs, dropping any edges with dangling references and flagging them as `invalid-reference` issues.**

The **Understand-Anything** engine from Egonex-AI builds a knowledge graph from codebases comprising **nodes**, **edges**, **layers**, and **tour steps**. When loading a graph, the system runs a multi-tiered validation pipeline to ensure structural integrity. This **graph schema validation** process guarantees that the final graph contains no dangling references, preventing downstream analysis and visualization failures.

## The Validation Pipeline in schema.ts

The core validation logic resides in [[`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts)](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/schema.ts). The `validateGraph` function orchestrates a pipeline that checks every component of the graph against defined schemas before the data enters the processing pipeline.

### Schema Compliance Checks

First, the validator ensures each edge conforms to `GraphEdgeSchema`. This verifies that the edge contains valid properties including `type`, `direction`, and `weight`. Only edges passing this structural check proceed to referential integrity verification.

### Source and Target Verification

The critical step for **referential integrity** occurs in the edge-validation block (lines 730-808). For every edge, the validator performs two essential lookups:

- **Source reference**: The `edge.source` identifier must exist in the set of valid node IDs.
- **Target reference**: The `edge.target` identifier must also be present in the node set.

If either identifier is missing, the validator records the failure and excludes the edge from the graph.

## How Referential Integrity Is Enforced

When the validator detects a missing node reference, it creates a validation issue with level `dropped` and category **`invalid-reference`** (lines 891-899). The specific error message identifies which edge and which property failed validation, such as `target "C" does not exist in nodes`. 

The system then filters out these invalid edges before returning the final graph in `result.data`. This guarantees that every remaining edge points to actual nodes, maintaining a consistent graph structure.

Beyond edges, the validator cleans up **layers** and **tour steps** by filtering out any `nodeIds` that no longer exist after node removal. This ensures the entire graph object remains internally consistent, not just the edge collection.

## Practical Example: Validating Graph References

The following TypeScript example demonstrates how dangling references are caught and removed:

```typescript
import { validateGraph } from "@understand-anything/core/schema";

const rawGraph = {
  version: "1.0.0",
  project: {
    name: "demo",
    languages: ["ts"],
    frameworks: [],
    description: "Demo graph",
    analyzedAt: new Date().toISOString(),
    gitCommitHash: "deadbeef",
  },
  nodes: [
    { id: "A", type: "file", name: "src/a.ts", summary: "", tags: [], complexity: "simple" },
    { id: "B", type: "function", name: "foo", summary: "", tags: [], complexity: "simple" },
  ],
  edges: [
    // ✅ valid edge – both nodes exist
    { source: "A", target: "B", type: "contains", direction: "forward", weight: 0.7 },
    // ❌ invalid edge – `target` does not exist
    { source: "A", target: "C", type: "contains", direction: "forward", weight: 0.5 },
  ],
  layers: [],
  tour: [],
};

const result = validateGraph(rawGraph);

if (!result.success) {
  console.error("Graph validation failed:", result.fatal);
}
console.log("Issues:", result.issues);
/*
  Issues will contain:
  {
    level: "dropped",
    category: "invalid-reference",
    message: 'edges[1]: target "C" does not exist in nodes — removed',
    path: 'edges[1].target'
  }
*/
console.log("Valid graph:", result.data);

```

In this example, the first edge remains in the graph because both node `A` and node `B` exist. The second edge is dropped because target `C` is not defined in the nodes array. The validation result preserves only the valid edge, ensuring the graph maintains strict referential integrity.

## Summary

- **Graph schema validation** in Understand-Anything uses the `validateGraph` function in [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts) to verify graph integrity.
- The system checks that every edge's `source` and `target` properties reference existing node IDs before accepting the graph.
- Invalid edges are categorized as **`invalid-reference`** issues with level `dropped` and excluded from the final graph (lines 891-899).
- The validator also cleans up layers and tour steps by removing references to non-existent nodes.
- This multi-tiered pipeline prevents dangling references from breaking downstream analyses or visualizations.

## Frequently Asked Questions

### What happens when an edge references a non-existent node?

The validator removes the edge from the final graph and adds an issue to the `result.issues` array with category `invalid-reference` and level `dropped`. The issue includes the specific path and message indicating which node ID was missing, allowing developers to trace the data problem.

### How does the validator handle layers and tour steps with invalid node references?

After validating edges, the system filters the `nodeIds` arrays within layers and tour steps to remove any identifiers that no longer exist in the validated node set. This ensures these collections remain consistent with the actual graph topology.

### Where is the validation logic implemented in the Understand-Anything codebase?

The primary validation logic resides in [[`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts)](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/schema.ts), specifically the `validateGraph` function. Edge referential integrity checks occur between lines 730-808, while issue logging happens at lines 891-899.

### Can validation failures be recovered or are dropped edges permanently removed?

Dropped edges are permanently excluded from `result.data` and cannot be automatically recovered. However, the validation result includes detailed issue records with paths and messages, allowing developers to fix the source data and re-run the validation process to include the corrected edges.