# KnowledgeGraph Schema in Understand-Anything: Complete Field Reference and Usage

> Explore the comprehensive KnowledgeGraph schema in Understand-Anything. Discover its fields for versioning, project details, nodes, edges, layers, and tour steps for in-depth code analysis.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: api-reference
- Published: 2026-06-19

---

**The KnowledgeGraph schema defines a typed, validated graph structure using Zod that includes version metadata, project information, nodes, edges, layers, and tour steps to represent codebases comprehensively.**

The KnowledgeGraph serves as the central data model in the [Understand-Anything](https://github.com/Egonex-AI/Understand-Anything) repository, representing entire codebases as structured, typed graphs. Defined in [`understand-anything-plugin/packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/schema.ts) using Zod schemas, it enforces type safety while providing automatic sanitization and auto-fixing capabilities for parsed code.

## Top-Level Structure of the KnowledgeGraph Schema

The root schema in [`schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/schema.ts) organizes codebase information into seven primary fields:

- **version**: Semantic version string of the graph format
- **kind**: Optional discriminator (`"codebase"` or `"knowledge"`) indicating the graph type
- **project**: Project-level metadata following `ProjectMetaSchema`
- **nodes**: Array of `GraphNodeSchema` representing entities like files, functions, and classes
- **edges**: Array of `GraphEdgeSchema` defining relationships between nodes
- **layers**: Array of `LayerSchema` for logical node groupings
- **tour**: Array of `TourStepSchema` containing pre-computed UI walkthrough steps

### Project Metadata (ProjectMetaSchema)

The `project` field stores repository-level information using `ProjectMetaSchema`, defined at lines 121-128 in [`schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/schema.ts):

```typescript
export const ProjectMetaSchema = z.object({
  name: z.string(),
  languages: z.array(z.string()),
  frameworks: z.array(z.string()),
  description: z.string(),
  analyzedAt: z.string(),
  gitCommitHash: z.string(),
});

```

This captures the **project name**, **programming languages**, **frameworks**, **description**, **analysis timestamp**, and **git commit hash** for complete traceability.

### Nodes and Entities (GraphNodeSchema)

The `nodes` array contains typed entities defined by `GraphNodeSchema` (lines 68-86). Each node requires:

- **id**: Unique string identifier
- **type**: Canonical node type (`"file"`, `"function"`, `"class"`, `"module"`, `"service"`, `"endpoint"`, `"article"`, etc.)
- **name**: Human-readable label
- **summary**: Short description of the entity
- **tags**: Array of classification strings
- **complexity**: Rating of `"simple"`, `"moderate"`, or `"complex"`

Optional fields include `filePath`, `lineRange`, `languageNotes`, `domainMeta`, and `knowledgeMeta` for extended metadata.

### Relationships and Edges (GraphEdgeSchema)

The `edges` array connects nodes via `GraphEdgeSchema` (lines 88-95), supporting 35 canonical edge types including `"calls"`, `"imports"`, `"depends_on"`, and `"deploys"`. Each edge contains:

- **source** and **target**: Node IDs defining the connection
- **type**: Relationship classification
- **direction**: `"forward"`, `"backward"`, or `"bidirectional"`
- **weight**: Numeric relevance score between 0 and 1
- Optional **description** for context

### Logical Groupings (LayerSchema)

The `layers` field enables architectural organization through `LayerSchema` (lines 97-102). Each layer groups related nodes with:

- **id**: Unique identifier
- **name**: Layer label (e.g., "UI layer", "Data layer")
- **description**: Purpose documentation
- **nodeIds**: Array of member node references

### Interactive Tours (TourStepSchema)

The `tour` field contains `TourStepSchema` entries (lines 104-110) that power UI walkthroughs:

- **order**: Sequence number
- **title** and **description**: Step content
- **nodeIds**: Nodes to highlight
- Optional **languageLesson**: Educational content

## Validation and Sanitization Functions

The [`schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/schema.ts) file exports four helper functions that ensure data integrity:

- **`validateGraph`**: Public entry point for schema validation
- **`sanitizeGraph`**: Removes invalid entries and maps LLM aliases to canonical types (e.g., `"func"` → `"function"`)
- **`autoFixGraph`**: Fills missing required fields with sensible defaults
- **`normalizeGraph`**: Coerces weight values to numbers and clamps them to `[0, 1]`

These functions return a detailed `GraphIssue` report when encountering invalid data, enabling comprehensive audit trails.

## Working with the KnowledgeGraph Schema

Import and validate graphs using the core schema utilities:

```typescript
import { KnowledgeGraphSchema, validateGraph } from "@understand-anything/core/schema";

// Load a raw JSON graph
const rawGraph = JSON.parse(await Deno.readTextFile("./my-graph.json"));

// Validate and auto-fix
const result = validateGraph(rawGraph);

if (result.success) {
  const graph = result.data!; // Typed as KnowledgeGraphSchema
  console.log("Graph version:", graph.version);
  console.log("Nodes:", graph.nodes.length);
  console.log("Edges:", graph.edges.length);
} else {
  console.error("Invalid graph:", result.errors);
  console.info("Issues:", result.issues);
}

```

Type-safe node filtering leverages TypeScript inference:

```typescript
import type { GraphNode } from "@understand-anything/core/schema";

function listFunctions(nodes: GraphNode[]) {
  return nodes
    .filter((n) => n.type === "function")
    .map((fn) => ({
      id: fn.id,
      name: fn.name,
      summary: fn.summary,
      complexity: fn.complexity,
    }));
}

const functions = listFunctions(result.data!.nodes);

```

## Summary

- The KnowledgeGraph schema in [`understand-anything-plugin/packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/schema.ts) defines a Zod-based type system for codebase representation
- **Seven top-level fields** capture version, project metadata, nodes, edges, layers, and tours
- **Node types** include files, functions, classes, services, and endpoints with complexity ratings
- **35 canonical edge types** model relationships like calls, imports, and dependencies with weighted relevance
- **Validation functions** (`validateGraph`, `sanitizeGraph`, `autoFixGraph`, `normalizeGraph`) ensure data integrity and auto-correction

## Frequently Asked Questions

### Where is the KnowledgeGraph schema defined?

The schema is defined in [`understand-anything-plugin/packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/schema.ts) in the Understand-Anything repository. This file contains all Zod schemas, type definitions, and validation logic for the graph structure.

### What types of nodes can be represented in the KnowledgeGraph?

The schema supports canonical node types including `"file"`, `"function"`, `"class"`, `"module"`, `"service"`, `"endpoint"`, and `"article"`. Each node includes required fields for ID, type, name, summary, tags, and complexity rating, with optional fields for file paths and line ranges.

### How does the schema handle invalid or incomplete data?

The `validateGraph` function along with `sanitizeGraph`, `autoFixGraph`, and `normalizeGraph` automatically map LLM aliases to canonical types, fill missing fields with defaults, coerce numeric weights to the range `[0, 1]`, and drop invalid entries while generating detailed `GraphIssue` reports for audit trails.

### What is the purpose of the tour field in the KnowledgeGraph?

The `tour` field contains pre-computed `TourStepSchema` arrays that define interactive UI walkthroughs. Each step specifies an order, title, description, and node IDs to highlight, enabling guided codebase exploration directly from the validated graph data.