How Understand-Anything Ensures Data Integrity in the Knowledge Graph Using Zod

Understand-Anything employs a six-stage tiered validation pipeline that uses Zod schemas to enforce strict type safety, completeness, and referential integrity across all knowledge graph operations.

Understand-Anything persists static analysis results in a structured knowledge graph comprised of nodes, edges, layers, and tour steps. To prevent runtime errors and ensure downstream components receive predictable data, the core package implements a comprehensive validation strategy. This system leverages data integrity in the knowledge graph using Zod through a sophisticated pipeline that combines preprocessing transformations with strict structural validation.

The Six-Stage Validation Pipeline

The validation logic resides in packages/core/src/schema.ts and orchestrates a multi-stage pipeline that progressively hardens graph data before it reaches application components.

1. Sanitize Raw Input

The sanitizeGraph function (lines 48-94) performs initial cleaning using pure JavaScript transformations. It normalises null and undefined values, lower-cases enum-like strings, and strips empty optional fields. This stage ensures that subsequent steps receive consistent primitive data types.

2. Normalise Aliases

LLM-generated outputs often contain variant terminology. The normalizeGraph function maps aliases like "func" to "function" or "extends" to "inherits" using centralized mapping objects: NODE_TYPE_ALIASES, EDGE_TYPE_ALIASES, COMPLEXITY_ALIASES, and DIRECTION_ALIASES (lines 16-46). This normalization guarantees that values match the canonical enumerations expected by Zod.

3. Auto-Fix Defaults and Coercions

Before schema validation, the autoFixGraph function (lines 96-166) injects missing required fields and coerces types. It converts string representations to numbers for fields like weight, clamps numeric ranges to valid bounds (e.g., z.number().min(0).max(1)), and supplies default arrays for missing tags. Each modification is recorded as a GraphIssue defined in packages/core/src/types.ts, maintaining full auditability.

4. Structural Validation with Zod

This critical stage applies strict Zod schemas to validate each collection and individual element. Fatal schema violations abort the process, while non-fatal issues cause specific elements to be dropped but reported. The schemas include:

  • GraphNodeSchema and GraphEdgeSchema (lines 68-71, 88-94)
  • LayerSchema and TourStepSchema
  • ProjectMetaSchema and KnowledgeGraphSchema (lines 21-29)

These definitions use z.object, z.enum, and z.array to enforce that every field conforms to explicit type contracts, including enumerations for node types and edge types (lines 4-14 and 70-76).

5. Referential Integrity Checks

After the Zod pass, validateGraph verifies cross-reference validity starting at line 73. It ensures every edge's source and target properties reference existing node IDs, and that layers and tour steps only contain valid node references. Invalid references are dropped and logged, preventing dangling pointers in the final graph.

6. Final Graph Assembly

The pipeline culminates in assembling a clean KnowledgeGraph object that conforms to KnowledgeGraphSchema. It injects a default version if missing and returns a ValidationResult containing the sanitized graph, the array of GraphIssue records, and a boolean success flag.

Zod Schema Architecture

The integrity guarantees rely on explicit schema definitions in packages/core/src/schema.ts. Enum definitions for EdgeTypeSchema and node type restrict values to supported categories, while field constraints enforce type safety through z.string(), z.number(), and chained validators like .min(0).max(1). The top-level KnowledgeGraphSchema validates the complete structure containing arrays of nodes, edges, layers, and metadata, ensuring the entire document conforms to the expected shape before downstream consumption.

Implementing Validation in Your Workflow

To validate a knowledge graph JSON file in your own implementation, import the validateGraph function from @understand-anything/core:

import { readFile } from "fs/promises";
import { validateGraph } from "@understand-anything/core";

async function loadAndValidate(path: string) {
  const raw = JSON.parse(await readFile(path, "utf8"));
  const result = validateGraph(raw);

  if (!result.success) {
    console.error("Graph validation failed:", result.errors);
    // result.issues contains detailed auto-fixed / dropped items
  } else {
    console.log("Validated graph version:", result.data?.version);
    // Use result.data safely – it conforms to KnowledgeGraphSchema
  }
}

loadAndValidate("./.understand-anything/knowledge-graph.json");

This function executes the full six-stage pipeline, returning a typed KnowledgeGraph that is safe for dashboard rendering and agent consumption.

Key Source Files

The validation system spans multiple modules in the Understand-Anything repository:

Summary

  • Understand-Anything ensures data integrity in the knowledge graph using Zod through a six-stage tiered validation pipeline that combines preprocessing with strict schema enforcement.
  • Pre-validation stages (sanitizeGraph, normalizeGraph, autoFixGraph) handle LLM inconsistencies, null values, and type coercion before Zod sees the data.
  • Zod schemas (GraphNodeSchema, GraphEdgeSchema, KnowledgeGraphSchema, etc.) in packages/core/src/schema.ts enforce type safety, enum constraints, and structural completeness.
  • Referential integrity checks verify that edges, layers, and tour steps reference valid node IDs after schema validation completes.
  • Every correction and dropped element is recorded as a GraphIssue, providing complete transparency into validation changes.

Frequently Asked Questions

What triggers a fatal validation error versus a dropped element?

Fatal errors occur when Zod encounters unrecoverable schema violations that autoFixGraph cannot resolve, such as completely missing required object fields or type mismatches that break the KnowledgeGraphSchema contract. In these cases, validateGraph returns success: false and aborts processing. Non-fatal issues—such as individual malformed nodes or edges that fail GraphNodeSchema validation—result in those specific elements being dropped while the rest of the graph processes successfully, with each drop recorded in the issues array.

How does the system handle LLM-generated type aliases?

The normalizeGraph function uses predefined alias maps (NODE_TYPE_ALIASES, EDGE_TYPE_ALIASES, COMPLEXITY_ALIASES, DIRECTION_ALIASES at lines 16-46 in schema.ts) to map variant strings like "func" to the canonical "function" or "extends" to "inherits". This normalization runs before the Zod validation stage, ensuring all values match the strict enumerations defined in EdgeTypeSchema and the node type enum (lines 4-14 and 70-76), preventing validation failures due to LLM terminology inconsistencies.

Can the validation pipeline be used independently of the analyzer?

Yes, the validateGraph function and associated schemas are exported from @understand-anything/core and can be imported into external projects. You can pass any JSON object representing a knowledge graph to validateGraph and receive a fully typed, validated KnowledgeGraph instance or detailed error reports. This makes the pipeline suitable for validating graphs generated by external tools or for sanitizing data before ingestion into other systems.

Where are default values defined for auto-fixed properties?

Default values are applied in the autoFixGraph function (lines 96-166 in packages/core/src/schema.ts). This function checks for missing type, complexity, and tags fields, injecting sensible defaults (such as empty arrays for tags or baseline complexity scores) and coercing numeric strings to proper numbers. Each auto-fix operation generates a GraphIssue record defined in packages/core/src/types.ts, allowing you to trace exactly which fields were modified during the validation process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →