KnowledgeGraph Schema in Understand-Anything: Complete Field Reference and Usage
The KnowledgeGraph schema defines a typed, validated graph structure using Zod that includes version metadata, project information, nodes, edges, layers, and tour steps to represent codebases comprehensively.
The KnowledgeGraph serves as the central data model in the Understand-Anything repository, representing entire codebases as structured, typed graphs. Defined in understand-anything-plugin/packages/core/src/schema.ts using Zod schemas, it enforces type safety while providing automatic sanitization and auto-fixing capabilities for parsed code.
Top-Level Structure of the KnowledgeGraph Schema
The root schema in schema.ts organizes codebase information into seven primary fields:
- version: Semantic version string of the graph format
- kind: Optional discriminator (
"codebase"or"knowledge") indicating the graph type - project: Project-level metadata following
ProjectMetaSchema - nodes: Array of
GraphNodeSchemarepresenting entities like files, functions, and classes - edges: Array of
GraphEdgeSchemadefining relationships between nodes - layers: Array of
LayerSchemafor logical node groupings - tour: Array of
TourStepSchemacontaining pre-computed UI walkthrough steps
Project Metadata (ProjectMetaSchema)
The project field stores repository-level information using ProjectMetaSchema, defined at lines 121-128 in schema.ts:
export const ProjectMetaSchema = z.object({
name: z.string(),
languages: z.array(z.string()),
frameworks: z.array(z.string()),
description: z.string(),
analyzedAt: z.string(),
gitCommitHash: z.string(),
});
This captures the project name, programming languages, frameworks, description, analysis timestamp, and git commit hash for complete traceability.
Nodes and Entities (GraphNodeSchema)
The nodes array contains typed entities defined by GraphNodeSchema (lines 68-86). Each node requires:
- id: Unique string identifier
- type: Canonical node type (
"file","function","class","module","service","endpoint","article", etc.) - name: Human-readable label
- summary: Short description of the entity
- tags: Array of classification strings
- complexity: Rating of
"simple","moderate", or"complex"
Optional fields include filePath, lineRange, languageNotes, domainMeta, and knowledgeMeta for extended metadata.
Relationships and Edges (GraphEdgeSchema)
The edges array connects nodes via GraphEdgeSchema (lines 88-95), supporting 35 canonical edge types including "calls", "imports", "depends_on", and "deploys". Each edge contains:
- source and target: Node IDs defining the connection
- type: Relationship classification
- direction:
"forward","backward", or"bidirectional" - weight: Numeric relevance score between 0 and 1
- Optional description for context
Logical Groupings (LayerSchema)
The layers field enables architectural organization through LayerSchema (lines 97-102). Each layer groups related nodes with:
- id: Unique identifier
- name: Layer label (e.g., "UI layer", "Data layer")
- description: Purpose documentation
- nodeIds: Array of member node references
Interactive Tours (TourStepSchema)
The tour field contains TourStepSchema entries (lines 104-110) that power UI walkthroughs:
- order: Sequence number
- title and description: Step content
- nodeIds: Nodes to highlight
- Optional languageLesson: Educational content
Validation and Sanitization Functions
The schema.ts file exports four helper functions that ensure data integrity:
validateGraph: Public entry point for schema validationsanitizeGraph: Removes invalid entries and maps LLM aliases to canonical types (e.g.,"func"→"function")autoFixGraph: Fills missing required fields with sensible defaultsnormalizeGraph: Coerces weight values to numbers and clamps them to[0, 1]
These functions return a detailed GraphIssue report when encountering invalid data, enabling comprehensive audit trails.
Working with the KnowledgeGraph Schema
Import and validate graphs using the core schema utilities:
import { KnowledgeGraphSchema, validateGraph } from "@understand-anything/core/schema";
// Load a raw JSON graph
const rawGraph = JSON.parse(await Deno.readTextFile("./my-graph.json"));
// Validate and auto-fix
const result = validateGraph(rawGraph);
if (result.success) {
const graph = result.data!; // Typed as KnowledgeGraphSchema
console.log("Graph version:", graph.version);
console.log("Nodes:", graph.nodes.length);
console.log("Edges:", graph.edges.length);
} else {
console.error("Invalid graph:", result.errors);
console.info("Issues:", result.issues);
}
Type-safe node filtering leverages TypeScript inference:
import type { GraphNode } from "@understand-anything/core/schema";
function listFunctions(nodes: GraphNode[]) {
return nodes
.filter((n) => n.type === "function")
.map((fn) => ({
id: fn.id,
name: fn.name,
summary: fn.summary,
complexity: fn.complexity,
}));
}
const functions = listFunctions(result.data!.nodes);
Summary
- The KnowledgeGraph schema in
understand-anything-plugin/packages/core/src/schema.tsdefines a Zod-based type system for codebase representation - Seven top-level fields capture version, project metadata, nodes, edges, layers, and tours
- Node types include files, functions, classes, services, and endpoints with complexity ratings
- 35 canonical edge types model relationships like calls, imports, and dependencies with weighted relevance
- Validation functions (
validateGraph,sanitizeGraph,autoFixGraph,normalizeGraph) ensure data integrity and auto-correction
Frequently Asked Questions
Where is the KnowledgeGraph schema defined?
The schema is defined in understand-anything-plugin/packages/core/src/schema.ts in the Understand-Anything repository. This file contains all Zod schemas, type definitions, and validation logic for the graph structure.
What types of nodes can be represented in the KnowledgeGraph?
The schema supports canonical node types including "file", "function", "class", "module", "service", "endpoint", and "article". Each node includes required fields for ID, type, name, summary, tags, and complexity rating, with optional fields for file paths and line ranges.
How does the schema handle invalid or incomplete data?
The validateGraph function along with sanitizeGraph, autoFixGraph, and normalizeGraph automatically map LLM aliases to canonical types, fill missing fields with defaults, coerce numeric weights to the range [0, 1], and drop invalid entries while generating detailed GraphIssue reports for audit trails.
What is the purpose of the tour field in the KnowledgeGraph?
The tour field contains pre-computed TourStepSchema arrays that define interactive UI walkthroughs. Each step specifies an order, title, description, and node IDs to highlight, enabling guided codebase exploration directly from the validated graph data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →