How Schema Validation Works When Loading a Knowledge Graph in Understand-Anything

Schema validation in Understand-Anything runs a four-tiered pipeline (sanitisation, normalisation, auto-fix, and strict Zod validation) when loading a graph, automatically correcting non-fatal issues while throwing fatal errors to prevent corrupted data from reaching the dashboard.

When working with the Egonex-AI/Understand-Anything repository, the loadGraph function enforces data integrity through a comprehensive validation system defined in packages/core/src/schema.ts. This pipeline ensures that only well-formed knowledge graphs reach the dashboard, providing detailed diagnostics about any corrections or removals necessary to sanitize the data.

The Four-Tier Validation Pipeline

The validation process operates sequentially through four distinct tiers, each handling specific data quality concerns before the graph reaches the UI.

Tier 1 – Sanitisation

The sanitizeGraph function (lines 48-94 in packages/core/src/schema.ts) normalises top-level collections and optional fields. It converts null collections into empty arrays, lower-cases enum-like strings, and strips null values from optional fields to create a consistent baseline for further processing.

Tier 2 – Normalisation

Next, normalizeGraph (lines 62-96) resolves LLM-generated aliases that may exist in the raw data. This function maps aliases for node-type, edge-type, complexity, and direction to canonical values—for example, transforming "func" into "function"—ensuring consistent terminology across the graph.

Tier 3 – Auto-Fix and Coercion

The autoFixGraph function (lines 96-150) fills missing fields with sensible defaults and coerces types where possible. It supplies defaults such as type: "file" and complexity: "moderate", clamps numeric weights to the range [0, 1], and maps remaining alias values. Each automatic correction is recorded as a GraphIssue with level auto-corrected, allowing developers to audit what changed.

Tier 4 – Strict Schema Validation

Finally, the system applies strict Zod schemas (GraphNodeSchema, GraphEdgeSchema, LayerSchema, TourStepSchema, and ProjectMetaSchema) to verify structural integrity. Invalid nodes, edges, layers, or tour steps are dropped from the graph and recorded as issues. This tier acts as the final gatekeeper before the data is considered valid.

Error Handling and Fatal Issues

The validation pipeline distinguishes between non-fatal and fatal problems through different handling mechanisms.

Non-Fatal Corrections

Missing fields, bad aliases, and out-of-range numeric values trigger automatic corrections. The validateGraph function returns a successful result containing the cleaned graph alongside an issues array describing what was fixed or dropped. The dashboard can display these warnings while still rendering the graph.

Fatal Validation Errors

Fatal problems abort the loading process entirely. According to the source code in packages/core/src/persistence/index.ts (lines 95-101), loadGraph throws an Error when encountering:

  • Missing or invalid project metadata (validated via ProjectMetaSchema)
  • Completely empty node lists
  • Malformed top-level collections that cannot be sanitised

The thrown exception propagates to the UI, which displays an error banner such as "Invalid knowledge graph: Missing or invalid project metadata" and prevents the corrupted data from rendering.

Implementation Details and Source Code

The validation system spans two critical files in the @understand-anything/core package:

  • packages/core/src/schema.ts: Contains the validateGraph function and all four pipeline tiers (sanitizeGraph, normalizeGraph, autoFixGraph), plus the Zod schema definitions.
  • packages/core/src/persistence/index.ts: Houses the loadGraph and loadDomainGraph entry points. When options?.validate !== false, it invokes validateGraph and throws on fatal errors.

The validation flow follows this exact sequence:

  1. loadGraph reads knowledge-graph.json (lines 85-94)
  2. Calls validateGraph(data) if validation is enabled (lines 99-102)
  3. Executes sanitisation, normalisation, and auto-fix sequentially
  4. Validates nodes with GraphNodeSchema (invalid nodes dropped)
  5. Validates edges with GraphEdgeSchema and checks referential integrity (lines 73-107)
  6. Filters dangling nodeIds from layers and tour steps
  7. Returns the clean graph object with the full issues list

Code Examples

Loading a Graph with Default Validation

import { loadGraph } from "@understand-anything/core";

const projectRoot = "/path/to/my/project";

try {
  const graph = loadGraph(projectRoot);  // validation enabled by default
  console.log("Graph loaded –", graph.nodes.length, "nodes");
} catch (e) {
  console.error("Failed to load graph:", e.message);
  // e.g. "Invalid knowledge graph: Missing or invalid project metadata"
}

Source: packages/core/src/persistence/index.ts (lines 85-102)

Loading Without Validation

const raw = loadGraph(projectRoot, { validate: false });
// `raw` may contain malformed data – use with caution for quick inspection only

Source: packages/core/src/persistence/index.ts (lines 94-103)

Inspecting Validation Issues Programmatically

import { validateGraph } from "@understand-anything/core";

const result = validateGraph(rawData);
if (!result.success) {
  console.error("Fatal error:", result.fatal);
}
console.log("All issues:", result.issues);

Source: packages/core/src/schema.ts (lines 99-102)

Example Auto-Fixed Output

{
  "version": "1.0.0",
  "project": { "name": "my-project", "root": "./src" },
  "nodes": [
    {
      "id": "file:src/index.ts",
      "type": "file",
      "name": "index.ts",
      "summary": "No summary",
      "tags": [],
      "complexity": "moderate"
    }
  ],
  "edges": [],
  "layers": [],
  "tour": []
}

This node was auto-corrected: missing summary defaulted to "No summary", missing tags became [], and missing complexity was set to "moderate" (see autoFixGraph lines 9-28).

Summary

  • Four-tier pipeline: Sanitisation → Normalisation → Auto-fix → Strict Zod validation, defined in packages/core/src/schema.ts.
  • Automatic correction: Non-fatal issues (missing fields, aliases, numeric ranges) are fixed and reported as GraphIssue entries with level auto-corrected.
  • Fatal error handling: Missing project metadata or empty node lists cause loadGraph to throw an Error, preventing the dashboard from rendering invalid data.
  • Entry points: Use loadGraph() from packages/core/src/persistence/index.ts with validation enabled by default, or pass { validate: false } to skip checks.
  • Referential integrity: Invalid nodes and edges are dropped during validation, with dangling references filtered from layers and tours.

Frequently Asked Questions

What happens if a node fails strict validation?

Invalid nodes are dropped from the graph during Tier 4 validation. The system records a GraphIssue for each removed node, but the loading process continues with the remaining valid nodes unless no valid nodes remain (which triggers a fatal error).

Can I disable schema validation when loading a graph?

Yes. Pass { validate: false } as the second argument to loadGraph() in packages/core/src/persistence/index.ts. This skips the entire validation pipeline and returns the raw data, though this is only recommended for quick inspection or debugging purposes.

How does the system handle LLM-generated aliases?

During the normalisation tier (normalizeGraph in packages/core/src/schema.ts), the system maps common aliases like "func" → "function" or "med" → "medium" for complexity levels. This ensures consistent terminology regardless of how the LLM labeled the data.

What constitutes a fatal validation error versus a fixable issue?

Fatal errors include missing project metadata (violating ProjectMetaSchema), completely empty node arrays, or unparseable top-level collections. These cause loadGraph to throw an exception. Fixable issues include missing optional fields, string aliases, or out-of-range numeric values, which autoFixGraph corrects automatically while reporting the changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →