# Egonex Understand-Anything Knowledge Graph JSON Schema Specification

> Discover the Egonex Understand-Anything knowledge graph JSON schema specification. Learn how repository structure becomes nodes and edges with Zod validation.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: api-reference
- Published: 2026-06-15

---

**The Egonex Understand-Anything knowledge graph follows a strict JSON schema defined in [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts), persisting repository structure as nodes and edges in [`./.understand-anything/knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/./.understand-anything/knowledge-graph.json) with Zod validation.**

The **Egonex Understand-Anything** tool constructs a semantic representation of code repositories as a structured knowledge graph. This graph is serialized to a JSON file that follows a rigorous schema specification, enabling interoperability between the analysis engine, web dashboard, and external tooling. The schema defines three core components—**nodes**, **edges**, and **metadata**—each with specific fields and type constraints enforced at runtime.

## Core Schema Structure

The knowledge graph JSON contains three top-level arrays and objects that define the entire repository structure.

### Top-Level Keys

- **`nodes`**: An array of `Node` objects representing every logical entity in the codebase (files, classes, functions, variables, tests)
- **`edges`**: An array of `Edge` objects defining relationships between nodes (imports, calls, containment)
- **`metadata`**: A `Metadata` object containing generation timestamps, tool version, and configuration settings

## Node Object Specification

Each entity in the repository becomes a node with standardized fields for identification and navigation.

**Required fields:**
- **`id`**: Global unique identifier (UUID-v4)
- **`type`**: Entity classification (`"file"`, `"module"`, `"class"`, `"function"`, `"variable"`, or other logical types)
- **`name`**: Human-readable identifier (e.g., `UserService`)
- **`path`**: Relative path from repository root (for file-type nodes)
- **`range`**: Source location object containing `start` and `end` positions with line/column coordinates
- **`language`**: Detected programming language identifier (`js`, `ts`, `py`, etc.)
- **`attributes`**: Open-ended key-value map (`Record<string, any>`) for extensible data like JSDoc comments or test status
- **`children`**: Array of node IDs representing direct descendants in the hierarchy
- **`parents`**: Array of node IDs enabling reverse navigation to ancestor nodes

## Edge Object Specification

Edges create the graph topology by linking related nodes with semantic relationship types.

**Edge properties:**
- **`source`**: UUID of the originating node
- **`target`**: UUID of the destination node
- **`type`**: Relationship classification including `"contains"`, `"calls"`, `"imports"`, `"extends"`, and `"references"`
- **`label`**: Optional human-readable description (e.g., `"uses"`)
- **`metadata`**: Optional map for edge-specific annotations

## Metadata Object Specification

The metadata section tracks provenance and configuration for reproducibility.

**Metadata fields:**
- **`generatedAt`**: ISO-8601 timestamp of graph creation
- **`toolVersion`**: Semantic version of Understand-Anything that produced the file
- **`repoRoot`**: Absolute filesystem path of the analyzed repository
- **`settings`**: Configuration object containing ignore patterns, active plugins, and generation options

## Schema Generation Pipeline

The knowledge graph JSON is produced through a four-stage pipeline implemented in the core package.

1. **File Discovery**: The tree-sitter plugin traverses the filesystem and parses each source file
2. **Extraction**: Language-specific extractors in `packages/core/src/plugins/extractors/*.ts` emit raw node data
3. **Normalization**: The [`graph-builder.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/graph-builder.ts) module consolidates raw nodes, resolves hierarchical relationships, and constructs edges
4. **Validation**: The Zod schema in [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts) validates the final object before serialization to [`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json)

## Working with the Schema

### Loading and Validating the Graph

Use the Zod schema exported from the core package to safely parse the JSON file:

```typescript
import { readFileSync } from "node:fs";
import { z } from "zod";
import { knowledgeGraphSchema } from "@understand-anything/core/schema";

const raw = readFileSync("./.understand-anything/knowledge-graph.json", "utf-8");
const graph = knowledgeGraphSchema.parse(JSON.parse(raw));

console.log(`Loaded ${graph.nodes.length} nodes and ${graph.edges.length} edges`);

```

### Traversing Node Relationships

Navigate from file nodes to their exported functions using the `children` array:

```typescript
function getExports(filePath: string) {
  const fileNode = graph.nodes.find(n => n.type === "file" && n.path === filePath);
  if (!fileNode) return [];

  return fileNode.children
    .map(id => graph.nodes.find(n => n.id === id))
    .filter(n => n && (n.type === "function" || n.type === "class"));
}

```

### Extending Nodes with Custom Attributes

Plugins can attach domain-specific data without schema modifications:

```typescript
// Inside a custom extractor
node.attributes = {
  ...node.attributes,
  myPluginScore: computeScore(node)
};

```

The open-ended `attributes` map accepts new keys automatically during Zod validation.

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts) | Zod schema defining the JSON shape and runtime validation |
| [`packages/core/src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/types.ts) | TypeScript interfaces mirroring the schema for compile-time safety |
| [`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts) | Core logic assembling raw nodes into a coherent graph |
| `packages/core/src/plugins/extractors/*.ts` | Language-specific extractors populating node objects |
| `scripts/generate-large-graph.mjs` | Utility for generating synthetic graphs for performance testing |

## Summary

- The **Egonex knowledge graph JSON schema** requires three top-level keys: `nodes`, `edges`, and `metadata`
- **Nodes** use UUID-v4 identifiers and support hierarchical navigation via `children` and `parents` arrays
- **Edges** define semantic relationships with constrained type enums like `"calls"` and `"imports"`
- The schema is enforced by **Zod** in [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts) before writing to [`./.understand-anything/knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/./.understand-anything/knowledge-graph.json)
- **Extensibility** is built-in through open-ended `attributes` and `metadata` maps that accept custom plugin data without breaking validation

## Frequently Asked Questions

### Where is the knowledge graph JSON file located by default?

By default, Understand-Anything writes the validated graph to [`./.understand-anything/knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/./.understand-anything/knowledge-graph.json) relative to the analyzed repository root. This path is configurable through the `settings` object stored in the metadata section.

### What validation library enforces the schema?

The schema uses **Zod** for runtime validation. The `knowledgeGraphSchema` object in [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts) validates the entire graph structure before serialization, ensuring that all nodes contain required fields like `id`, `type`, and `name`, and that edge `source` and `target` properties reference valid node IDs.

### Can I add custom fields to nodes without modifying the core schema?

Yes. The `attributes` field on nodes and the `metadata` field on edges are typed as `Record<string, any>`, allowing plugins to attach arbitrary data. The Zod validation accepts these open-ended maps, so custom keys like `myPluginScore` or `testCoverage` persist without requiring changes to [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts).

### What are the valid node types in the Understand-Anything schema?

The schema supports a flexible enumeration including `"file"`, `"module"`, `"class"`, `"function"`, `"variable"`, and additional language-specific types. The `type` field is validated as a string, with specific extractors in `packages/core/src/plugins/extractors/*.ts` determining the appropriate classification based on source code analysis.