# How the Knowledge Graph Schema Is Validated Upon Loading in Egonex-AI Understand-Anything

> Discover how Egonex-AI validates knowledge graph schema loading through a four-tier pipeline that sanitizes, normalizes, auto-fixes, and enforces Zod constraints for robust data integrity.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: internals
- Published: 2026-06-09

---

**The knowledge graph schema is validated through a four-tier pipeline in [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts) that sanitizes input, normalizes LLM aliases, auto-fixes missing fields, and enforces strict Zod schema constraints, returning either a clean graph with warnings or a fatal error for critical structural failures.**

When loading a knowledge graph JSON file in the Egonex-AI Understand-Anything project, the core validation engine ensures data integrity through a rigorous multi-stage process. This validation pipeline, implemented in the `@understand-anything/core` package, automatically sanitizes LLM-generated artifacts, normalizes type aliases, and enforces strict schema constraints before the data reaches the dashboard UI.

## The Four-Stage Validation Pipeline

The validation process defined in [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts) executes four distinct tiers that progressively refine the raw graph data into a strictly typed, validated structure.

### Stage 1: Sanitization with `sanitizeGraph`

The pipeline begins with `sanitizeGraph`, which normalizes null values, lower-cases enum-like strings, and removes empty optional fields. This function handles lines 48-94 of [`schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/schema.ts), ensuring that inconsistently formatted data from external sources enters the pipeline in a predictable state.

### Stage 2: Alias Normalization with `normalizeGraph`

Next, `normalizeGraph` (lines 68-80) replaces LLM-generated type and edge aliases with canonical enum values. For example, the system converts `"func"` to `"function"` and other shorthand identifiers to their standardized counterparts, ensuring consistency across graphs generated by different language models.

### Stage 3: Auto-Fixing and Coercion with `autoFixGraph`

The `autoFixGraph` function (lines 96-150) supplies sensible defaults for missing fields and coerces types where possible. When a node lacks a type definition, it defaults to `"file"`; when an edge lacks a weight, it defaults to `0.5`. The function also handles type coercion, such as converting string representations to numbers for the `weight` field, while accumulating a list of non-fatal issues for reporting.

### Stage 4: Strict Validation with `validateGraph`

The `validateGraph` function orchestrates the complete pipeline and performs final strict validation. After executing sanitization, normalization, and auto-fixing, it:

- Verifies that top-level collections are arrays (fatal error if not) at lines 118-121
- Validates required project metadata against `ProjectMetaSchema` (lines 332-338)
- Parses each node with `GraphNodeSchema`, dropping invalid nodes and recording them in `issues` (lines 443-456)
- Validates edges with `GraphEdgeSchema`, ensuring `source` and `target` IDs exist in the validated node set, dropping broken edges (lines 470-506)
- Validates layers and tour steps, stripping dangling references (lines 511-540)

## Fatal Errors vs. Auto-Correction

The validation pipeline distinguishes between recoverable issues and critical structural failures. **Fatal errors** occur when the root object is not an object, when required project metadata is missing, or when no valid nodes remain after parsing. In these cases, `validateGraph` returns `success: false` with a `fatal` message explaining the constraint violation.

**Non-fatal issues** include auto-corrected type coercions, missing default values, and dropped invalid nodes or edges. These accumulate in the `issues` array while returning `success: true` with the sanitized graph data, allowing the dashboard to display warnings without blocking functionality.

## Implementation Example

To validate a knowledge graph JSON file in your own implementation:

```typescript
import { validateGraph } from "@understand-anything/core";

// Load a JSON file that was generated by the analyzer
const rawGraph = await fetch("/file-content.json").then(r => r.json());

// Run the full validation pipeline
const result = validateGraph(rawGraph);

if (result.success) {
  console.log("✅ Graph is valid");
  // Use the clean data – e.g. feed it to the dashboard
  const graph = result.data!;
} else {
  console.error("❌ Graph validation failed:", result.fatal);
  console.warn("Issues discovered:", result.issues);
}

```

The same function is called by the dashboard server when loading a project's [`.understand-anything/knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json) file, ensuring that the UI only receives well-typed graph data.

## Key Source Files

- **[`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts)** – Contains all Zod schemas, sanitization, alias normalization, auto-fixing, and the top-level `validateGraph` function
- **[`packages/core/src/__tests__/schema.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/__tests__/schema.test.ts)** – Test suite exercising the validation pipeline with edge cases, auto-fix scenarios, and invalid entry handling
- **[`packages/core/src/languages/configs/json-schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/languages/configs/json-schema.ts)** – Provides JSON-Schema support for the graph when exporting to external tools

## Summary

- The validation pipeline runs in four sequential stages: sanitization, alias normalization, auto-fixing, and strict validation
- **Invalid nodes are dropped** and recorded in the issues array rather than causing complete failure
- **Missing edge weights default to 0.5** and missing node types default to `"file"` through the auto-fixing layer
- **LLM-generated aliases** like `"func"` are normalized to canonical values such as `"function"`
- The `validateGraph` function in [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts) returns a discriminated union with `success: true` and clean data, or `success: false` with a fatal error message

## Frequently Asked Questions

### What happens if a node fails validation during the strict validation stage?

If a node fails to parse against `GraphNodeSchema`, the validator drops that specific node from the graph and records the failure in the `issues` array. The validation continues processing remaining nodes and edges, allowing the graph to load with partial data rather than failing completely.

### How does the pipeline handle LLM-generated aliases in the knowledge graph?

The `normalizeGraph` function (lines 68-80) replaces LLM-generated shorthand aliases with canonical enum values. For example, it converts `"func"` to `"function"` and other non-standard type identifiers to their standardized counterparts, ensuring consistency across different model outputs.

### What default values are applied when fields are missing from the graph?

The `autoFixGraph` function supplies `"file"` as the default node type and `0.5` as the default edge weight when these fields are missing. It also coerces string values to numbers where appropriate, such as converting string representations of weights to numeric values.

### Where is the knowledge graph schema validation logic tested?

The validation logic is comprehensively tested in [`packages/core/src/__tests__/schema.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/__tests__/schema.test.ts), which exercises edge cases, auto-fix scenarios, and the handling of invalid entries to ensure the pipeline robustly handles malformed input from various LLM-generated sources.