KnowledgeGraph Schema in Understand-Anything: Complete Field Reference and Usage

The KnowledgeGraph schema defines a typed, validated graph structure using Zod that includes version metadata, project information, nodes, edges, layers, and tour steps to represent codebases comprehensively.

The KnowledgeGraph serves as the central data model in the Understand-Anything repository, representing entire codebases as structured, typed graphs. Defined in understand-anything-plugin/packages/core/src/schema.ts using Zod schemas, it enforces type safety while providing automatic sanitization and auto-fixing capabilities for parsed code.

Top-Level Structure of the KnowledgeGraph Schema

The root schema in schema.ts organizes codebase information into seven primary fields:

  • version: Semantic version string of the graph format
  • kind: Optional discriminator ("codebase" or "knowledge") indicating the graph type
  • project: Project-level metadata following ProjectMetaSchema
  • nodes: Array of GraphNodeSchema representing entities like files, functions, and classes
  • edges: Array of GraphEdgeSchema defining relationships between nodes
  • layers: Array of LayerSchema for logical node groupings
  • tour: Array of TourStepSchema containing pre-computed UI walkthrough steps

Project Metadata (ProjectMetaSchema)

The project field stores repository-level information using ProjectMetaSchema, defined at lines 121-128 in schema.ts:

export const ProjectMetaSchema = z.object({
  name: z.string(),
  languages: z.array(z.string()),
  frameworks: z.array(z.string()),
  description: z.string(),
  analyzedAt: z.string(),
  gitCommitHash: z.string(),
});

This captures the project name, programming languages, frameworks, description, analysis timestamp, and git commit hash for complete traceability.

Nodes and Entities (GraphNodeSchema)

The nodes array contains typed entities defined by GraphNodeSchema (lines 68-86). Each node requires:

  • id: Unique string identifier
  • type: Canonical node type ("file", "function", "class", "module", "service", "endpoint", "article", etc.)
  • name: Human-readable label
  • summary: Short description of the entity
  • tags: Array of classification strings
  • complexity: Rating of "simple", "moderate", or "complex"

Optional fields include filePath, lineRange, languageNotes, domainMeta, and knowledgeMeta for extended metadata.

Relationships and Edges (GraphEdgeSchema)

The edges array connects nodes via GraphEdgeSchema (lines 88-95), supporting 35 canonical edge types including "calls", "imports", "depends_on", and "deploys". Each edge contains:

  • source and target: Node IDs defining the connection
  • type: Relationship classification
  • direction: "forward", "backward", or "bidirectional"
  • weight: Numeric relevance score between 0 and 1
  • Optional description for context

Logical Groupings (LayerSchema)

The layers field enables architectural organization through LayerSchema (lines 97-102). Each layer groups related nodes with:

  • id: Unique identifier
  • name: Layer label (e.g., "UI layer", "Data layer")
  • description: Purpose documentation
  • nodeIds: Array of member node references

Interactive Tours (TourStepSchema)

The tour field contains TourStepSchema entries (lines 104-110) that power UI walkthroughs:

  • order: Sequence number
  • title and description: Step content
  • nodeIds: Nodes to highlight
  • Optional languageLesson: Educational content

Validation and Sanitization Functions

The schema.ts file exports four helper functions that ensure data integrity:

  • validateGraph: Public entry point for schema validation
  • sanitizeGraph: Removes invalid entries and maps LLM aliases to canonical types (e.g., "func" → "function")
  • autoFixGraph: Fills missing required fields with sensible defaults
  • normalizeGraph: Coerces weight values to numbers and clamps them to [0, 1]

These functions return a detailed GraphIssue report when encountering invalid data, enabling comprehensive audit trails.

Working with the KnowledgeGraph Schema

Import and validate graphs using the core schema utilities:

import { KnowledgeGraphSchema, validateGraph } from "@understand-anything/core/schema";

// Load a raw JSON graph
const rawGraph = JSON.parse(await Deno.readTextFile("./my-graph.json"));

// Validate and auto-fix
const result = validateGraph(rawGraph);

if (result.success) {
  const graph = result.data!; // Typed as KnowledgeGraphSchema
  console.log("Graph version:", graph.version);
  console.log("Nodes:", graph.nodes.length);
  console.log("Edges:", graph.edges.length);
} else {
  console.error("Invalid graph:", result.errors);
  console.info("Issues:", result.issues);
}

Type-safe node filtering leverages TypeScript inference:

import type { GraphNode } from "@understand-anything/core/schema";

function listFunctions(nodes: GraphNode[]) {
  return nodes
    .filter((n) => n.type === "function")
    .map((fn) => ({
      id: fn.id,
      name: fn.name,
      summary: fn.summary,
      complexity: fn.complexity,
    }));
}

const functions = listFunctions(result.data!.nodes);

Summary

  • The KnowledgeGraph schema in understand-anything-plugin/packages/core/src/schema.ts defines a Zod-based type system for codebase representation
  • Seven top-level fields capture version, project metadata, nodes, edges, layers, and tours
  • Node types include files, functions, classes, services, and endpoints with complexity ratings
  • 35 canonical edge types model relationships like calls, imports, and dependencies with weighted relevance
  • Validation functions (validateGraph, sanitizeGraph, autoFixGraph, normalizeGraph) ensure data integrity and auto-correction

Frequently Asked Questions

Where is the KnowledgeGraph schema defined?

The schema is defined in understand-anything-plugin/packages/core/src/schema.ts in the Understand-Anything repository. This file contains all Zod schemas, type definitions, and validation logic for the graph structure.

What types of nodes can be represented in the KnowledgeGraph?

The schema supports canonical node types including "file", "function", "class", "module", "service", "endpoint", and "article". Each node includes required fields for ID, type, name, summary, tags, and complexity rating, with optional fields for file paths and line ranges.

How does the schema handle invalid or incomplete data?

The validateGraph function along with sanitizeGraph, autoFixGraph, and normalizeGraph automatically map LLM aliases to canonical types, fill missing fields with defaults, coerce numeric weights to the range [0, 1], and drop invalid entries while generating detailed GraphIssue reports for audit trails.

What is the purpose of the tour field in the KnowledgeGraph?

The tour field contains pre-computed TourStepSchema arrays that define interactive UI walkthroughs. Each step specifies an order, title, description, and node IDs to highlight, enabling guided codebase exploration directly from the validated graph data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →