How the Knowledge Graph Schema Is Validated Upon Loading in Egonex-AI Understand-Anything
The knowledge graph schema is validated through a four-tier pipeline in packages/core/src/schema.ts that sanitizes input, normalizes LLM aliases, auto-fixes missing fields, and enforces strict Zod schema constraints, returning either a clean graph with warnings or a fatal error for critical structural failures.
When loading a knowledge graph JSON file in the Egonex-AI Understand-Anything project, the core validation engine ensures data integrity through a rigorous multi-stage process. This validation pipeline, implemented in the @understand-anything/core package, automatically sanitizes LLM-generated artifacts, normalizes type aliases, and enforces strict schema constraints before the data reaches the dashboard UI.
The Four-Stage Validation Pipeline
The validation process defined in packages/core/src/schema.ts executes four distinct tiers that progressively refine the raw graph data into a strictly typed, validated structure.
Stage 1: Sanitization with sanitizeGraph
The pipeline begins with sanitizeGraph, which normalizes null values, lower-cases enum-like strings, and removes empty optional fields. This function handles lines 48-94 of schema.ts, ensuring that inconsistently formatted data from external sources enters the pipeline in a predictable state.
Stage 2: Alias Normalization with normalizeGraph
Next, normalizeGraph (lines 68-80) replaces LLM-generated type and edge aliases with canonical enum values. For example, the system converts "func" to "function" and other shorthand identifiers to their standardized counterparts, ensuring consistency across graphs generated by different language models.
Stage 3: Auto-Fixing and Coercion with autoFixGraph
The autoFixGraph function (lines 96-150) supplies sensible defaults for missing fields and coerces types where possible. When a node lacks a type definition, it defaults to "file"; when an edge lacks a weight, it defaults to 0.5. The function also handles type coercion, such as converting string representations to numbers for the weight field, while accumulating a list of non-fatal issues for reporting.
Stage 4: Strict Validation with validateGraph
The validateGraph function orchestrates the complete pipeline and performs final strict validation. After executing sanitization, normalization, and auto-fixing, it:
- Verifies that top-level collections are arrays (fatal error if not) at lines 118-121
- Validates required project metadata against
ProjectMetaSchema(lines 332-338) - Parses each node with
GraphNodeSchema, dropping invalid nodes and recording them inissues(lines 443-456) - Validates edges with
GraphEdgeSchema, ensuringsourceandtargetIDs exist in the validated node set, dropping broken edges (lines 470-506) - Validates layers and tour steps, stripping dangling references (lines 511-540)
Fatal Errors vs. Auto-Correction
The validation pipeline distinguishes between recoverable issues and critical structural failures. Fatal errors occur when the root object is not an object, when required project metadata is missing, or when no valid nodes remain after parsing. In these cases, validateGraph returns success: false with a fatal message explaining the constraint violation.
Non-fatal issues include auto-corrected type coercions, missing default values, and dropped invalid nodes or edges. These accumulate in the issues array while returning success: true with the sanitized graph data, allowing the dashboard to display warnings without blocking functionality.
Implementation Example
To validate a knowledge graph JSON file in your own implementation:
import { validateGraph } from "@understand-anything/core";
// Load a JSON file that was generated by the analyzer
const rawGraph = await fetch("/file-content.json").then(r => r.json());
// Run the full validation pipeline
const result = validateGraph(rawGraph);
if (result.success) {
console.log("✅ Graph is valid");
// Use the clean data – e.g. feed it to the dashboard
const graph = result.data!;
} else {
console.error("❌ Graph validation failed:", result.fatal);
console.warn("Issues discovered:", result.issues);
}
The same function is called by the dashboard server when loading a project's .understand-anything/knowledge-graph.json file, ensuring that the UI only receives well-typed graph data.
Key Source Files
packages/core/src/schema.ts– Contains all Zod schemas, sanitization, alias normalization, auto-fixing, and the top-levelvalidateGraphfunctionpackages/core/src/__tests__/schema.test.ts– Test suite exercising the validation pipeline with edge cases, auto-fix scenarios, and invalid entry handlingpackages/core/src/languages/configs/json-schema.ts– Provides JSON-Schema support for the graph when exporting to external tools
Summary
- The validation pipeline runs in four sequential stages: sanitization, alias normalization, auto-fixing, and strict validation
- Invalid nodes are dropped and recorded in the issues array rather than causing complete failure
- Missing edge weights default to 0.5 and missing node types default to
"file"through the auto-fixing layer - LLM-generated aliases like
"func"are normalized to canonical values such as"function" - The
validateGraphfunction inpackages/core/src/schema.tsreturns a discriminated union withsuccess: trueand clean data, orsuccess: falsewith a fatal error message
Frequently Asked Questions
What happens if a node fails validation during the strict validation stage?
If a node fails to parse against GraphNodeSchema, the validator drops that specific node from the graph and records the failure in the issues array. The validation continues processing remaining nodes and edges, allowing the graph to load with partial data rather than failing completely.
How does the pipeline handle LLM-generated aliases in the knowledge graph?
The normalizeGraph function (lines 68-80) replaces LLM-generated shorthand aliases with canonical enum values. For example, it converts "func" to "function" and other non-standard type identifiers to their standardized counterparts, ensuring consistency across different model outputs.
What default values are applied when fields are missing from the graph?
The autoFixGraph function supplies "file" as the default node type and 0.5 as the default edge weight when these fields are missing. It also coerces string values to numbers where appropriate, such as converting string representations of weights to numeric values.
Where is the knowledge graph schema validation logic tested?
The validation logic is comprehensively tested in packages/core/src/__tests__/schema.test.ts, which exercises edge cases, auto-fix scenarios, and the handling of invalid entries to ensure the pipeline robustly handles malformed input from various LLM-generated sources.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →