Egonex Understand-Anything Knowledge Graph JSON Schema Specification
The Egonex Understand-Anything knowledge graph follows a strict JSON schema defined in packages/core/src/schema.ts, persisting repository structure as nodes and edges in ./.understand-anything/knowledge-graph.json with Zod validation.
The Egonex Understand-Anything tool constructs a semantic representation of code repositories as a structured knowledge graph. This graph is serialized to a JSON file that follows a rigorous schema specification, enabling interoperability between the analysis engine, web dashboard, and external tooling. The schema defines three core components—nodes, edges, and metadata—each with specific fields and type constraints enforced at runtime.
Core Schema Structure
The knowledge graph JSON contains three top-level arrays and objects that define the entire repository structure.
Top-Level Keys
nodes: An array ofNodeobjects representing every logical entity in the codebase (files, classes, functions, variables, tests)edges: An array ofEdgeobjects defining relationships between nodes (imports, calls, containment)metadata: AMetadataobject containing generation timestamps, tool version, and configuration settings
Node Object Specification
Each entity in the repository becomes a node with standardized fields for identification and navigation.
Required fields:
id: Global unique identifier (UUID-v4)type: Entity classification ("file","module","class","function","variable", or other logical types)name: Human-readable identifier (e.g.,UserService)path: Relative path from repository root (for file-type nodes)range: Source location object containingstartandendpositions with line/column coordinateslanguage: Detected programming language identifier (js,ts,py, etc.)attributes: Open-ended key-value map (Record<string, any>) for extensible data like JSDoc comments or test statuschildren: Array of node IDs representing direct descendants in the hierarchyparents: Array of node IDs enabling reverse navigation to ancestor nodes
Edge Object Specification
Edges create the graph topology by linking related nodes with semantic relationship types.
Edge properties:
source: UUID of the originating nodetarget: UUID of the destination nodetype: Relationship classification including"contains","calls","imports","extends", and"references"label: Optional human-readable description (e.g.,"uses")metadata: Optional map for edge-specific annotations
Metadata Object Specification
The metadata section tracks provenance and configuration for reproducibility.
Metadata fields:
generatedAt: ISO-8601 timestamp of graph creationtoolVersion: Semantic version of Understand-Anything that produced the filerepoRoot: Absolute filesystem path of the analyzed repositorysettings: Configuration object containing ignore patterns, active plugins, and generation options
Schema Generation Pipeline
The knowledge graph JSON is produced through a four-stage pipeline implemented in the core package.
- File Discovery: The tree-sitter plugin traverses the filesystem and parses each source file
- Extraction: Language-specific extractors in
packages/core/src/plugins/extractors/*.tsemit raw node data - Normalization: The
graph-builder.tsmodule consolidates raw nodes, resolves hierarchical relationships, and constructs edges - Validation: The Zod schema in
packages/core/src/schema.tsvalidates the final object before serialization toknowledge-graph.json
Working with the Schema
Loading and Validating the Graph
Use the Zod schema exported from the core package to safely parse the JSON file:
import { readFileSync } from "node:fs";
import { z } from "zod";
import { knowledgeGraphSchema } from "@understand-anything/core/schema";
const raw = readFileSync("./.understand-anything/knowledge-graph.json", "utf-8");
const graph = knowledgeGraphSchema.parse(JSON.parse(raw));
console.log(`Loaded ${graph.nodes.length} nodes and ${graph.edges.length} edges`);
Traversing Node Relationships
Navigate from file nodes to their exported functions using the children array:
function getExports(filePath: string) {
const fileNode = graph.nodes.find(n => n.type === "file" && n.path === filePath);
if (!fileNode) return [];
return fileNode.children
.map(id => graph.nodes.find(n => n.id === id))
.filter(n => n && (n.type === "function" || n.type === "class"));
}
Extending Nodes with Custom Attributes
Plugins can attach domain-specific data without schema modifications:
// Inside a custom extractor
node.attributes = {
...node.attributes,
myPluginScore: computeScore(node)
};
The open-ended attributes map accepts new keys automatically during Zod validation.
Key Implementation Files
| File | Purpose |
|---|---|
packages/core/src/schema.ts |
Zod schema defining the JSON shape and runtime validation |
packages/core/src/types.ts |
TypeScript interfaces mirroring the schema for compile-time safety |
packages/core/src/analyzer/graph-builder.ts |
Core logic assembling raw nodes into a coherent graph |
packages/core/src/plugins/extractors/*.ts |
Language-specific extractors populating node objects |
scripts/generate-large-graph.mjs |
Utility for generating synthetic graphs for performance testing |
Summary
- The Egonex knowledge graph JSON schema requires three top-level keys:
nodes,edges, andmetadata - Nodes use UUID-v4 identifiers and support hierarchical navigation via
childrenandparentsarrays - Edges define semantic relationships with constrained type enums like
"calls"and"imports" - The schema is enforced by Zod in
packages/core/src/schema.tsbefore writing to./.understand-anything/knowledge-graph.json - Extensibility is built-in through open-ended
attributesandmetadatamaps that accept custom plugin data without breaking validation
Frequently Asked Questions
Where is the knowledge graph JSON file located by default?
By default, Understand-Anything writes the validated graph to ./.understand-anything/knowledge-graph.json relative to the analyzed repository root. This path is configurable through the settings object stored in the metadata section.
What validation library enforces the schema?
The schema uses Zod for runtime validation. The knowledgeGraphSchema object in packages/core/src/schema.ts validates the entire graph structure before serialization, ensuring that all nodes contain required fields like id, type, and name, and that edge source and target properties reference valid node IDs.
Can I add custom fields to nodes without modifying the core schema?
Yes. The attributes field on nodes and the metadata field on edges are typed as Record<string, any>, allowing plugins to attach arbitrary data. The Zod validation accepts these open-ended maps, so custom keys like myPluginScore or testCoverage persist without requiring changes to packages/core/src/schema.ts.
What are the valid node types in the Understand-Anything schema?
The schema supports a flexible enumeration including "file", "module", "class", "function", "variable", and additional language-specific types. The type field is validated as a string, with specific extractors in packages/core/src/plugins/extractors/*.ts determining the appropriate classification based on source code analysis.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →