How Understand Anything Builds Deterministic Structural Graphs with Semantic Intent Using Tree‑Sitter and LLMs
Understand Anything generates deterministic knowledge graphs by using Tree‑Sitter to produce repeatable AST structures and LLMs to decorate nodes with semantic metadata, ensuring identical source code always yields identical graph topology while enriching entities with human‑readable intent.
The Egonex‑AI/Understand‑Anything repository solves codebase comprehension by strictly separating structural analysis from semantic interpretation. This architecture guarantees deterministic structural graphs with semantic intent—where the graph topology never changes between runs unless the source changes, while LLM‑generated summaries provide searchable, human‑friendly context.
The Hybrid Architecture: Deterministic Structure + LLM Intent
The system rests on two orthogonal techniques. First, Tree‑Sitter provides pure‑static, grammar‑driven parsing that yields a repeatable abstract syntax tree (AST) for every supported file. Second, LLM‑driven semantic enrichment decorates these structural nodes with summaries, tags, and complexity scores via prompt‑based analysis.
This separation ensures that the structural graph is fully deterministic and cacheable, while the semantic layer adds volatile but non‑structural metadata.
The Six‑Step Pipeline
Understand Anything processes each file through a strict pipeline defined in the core analyzer.
Step 1: Language Detection
The LanguageRegistry maps file paths to language IDs based on extensions. Located in understand-anything-plugin/packages/core/src/languages/language-registry.ts, this registry feeds the correct grammar to the parser.
Step 2: Deterministic Parsing
TreeSitterPlugin.init() loads WebAssembly grammars once during initialization. Calling TreeSitterPlugin.getParser(filePath) returns a pre‑configured parser, and parser.parse(content) yields the same AST on every execution. This logic resides in understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts.
Step 3: Structural Extraction
Each language implements a pure extractor (e.g., typescript-extractor) that walks the AST and returns a StructuralAnalysis object containing functions, classes, imports, and exports. These extractors contain no randomness or external I/O. They are aggregated in understand-anything-plugin/packages/core/src/plugins/extractors/index.ts.
Step 4: Graph Construction
GraphBuilder.addFileWithAnalysis() in understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts creates file nodes and deterministic child nodes using stable IDs like function:src/app.ts:myFunc. Edges (contains, imports, calls) are added only once using a deduplication set (edgeKeys), guaranteeing identical source produces identical topology.
Step 5: Semantic Intent Injection
After the structural graph exists, the LLM analyzer enriches nodes. The buildFileAnalysisPrompt and buildProjectSummaryPrompt functions in understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts feed raw source and project context to the model. Responses are parsed via parseFileAnalysisResponse and parseProjectSummaryResponse, then injected as summary, tags, and complexity metadata.
Step 6: Final Knowledge Graph Assembly
GraphBuilder.build() bundles the deterministic node and edge collections with LLM‑populated metadata into a KnowledgeGraph object.
Implementation Deep Dive
Key Source Files
The deterministic pipeline relies on five critical files:
tree-sitter-plugin.ts: Loads WASM grammars and delegates to language‑specific extractors.graph-builder.ts: Constructs deterministic nodes/edges and manages theedgeKeysdeduplication set.llm-analyzer.ts: Generates prompts and parses LLM responses for semantic enrichment.language-registry.ts: Detects languages from file extensions.extractors/index.ts: Houses pure functions that walk the Tree‑sitter AST.
End‑to‑End Code Example
The following TypeScript script demonstrates the complete pipeline:
import { TreeSitterPlugin } from "./packages/core/src/plugins/tree-sitter-plugin.js";
import { GraphBuilder } from "./packages/core/src/analyzer/graph-builder.js";
import { LanguageRegistry } from "./packages/core/src/languages/language-registry.js";
import {
buildFileAnalysisPrompt,
parseFileAnalysisResponse,
} from "./packages/core/src/analyzer/llm-analyzer.js";
// 1️⃣ Initialise the tree‑sitter plugin with the default language configs
const tsPlugin = new TreeSitterPlugin();
await tsPlugin.init(); // <-- loads WASM grammars only once
// 2️⃣ Prepare a GraphBuilder for the project
const registry = LanguageRegistry.createDefault();
const graph = new GraphBuilder("my‑project", "deadbeef", registry);
// 3️⃣ Analyse a source file
const filePath = "src/example.ts";
const content = await Deno.readTextFile(filePath); // or fs.readFileSync(...)
const structural = tsPlugin.analyzeFile(filePath, content);
// 4️⃣ Enrich with LLM intent (pseudo‑LLM call)
const projectContext = "A simple TypeScript CLI tool";
const prompt = buildFileAnalysisPrompt(filePath, content, projectContext);
// `llmResponse` would be the raw string from the model
// const llmResponse = await callYourLLM(prompt);
const llmResponse = `{
"fileSummary":"CLI parses arguments and prints a greeting",
"tags":["cli","utility"],
"complexity":"simple",
"functionSummaries":{"main":"Entry point that wires everything together"},
"classSummaries":{}
}`;
const llmMeta = parseFileAnalysisResponse(llmResponse)!;
// 5️⃣ Add the file plus its analysis to the graph
graph.addFileWithAnalysis(filePath, structural, {
fileSummary: llmMeta.fileSummary,
summaries: llmMeta.functionSummaries,
tags: llmMeta.tags,
complexity: llmMeta.complexity,
});
// 6️⃣ Build the final deterministic graph
const knowledgeGraph = graph.build();
console.log(JSON.stringify(knowledgeGraph, null, 2));
Running this script twice on identical source produces identical JSON output for node IDs and edge topology. Only the LLM‑generated text fields vary, and these are stored as metadata without affecting graph structure.
Why Determinism Matters
Deterministic structural graphs enable reliable caching, diff detection, and version control integration. Because the AST extraction and graph construction contain no randomness—verified by the pure functions in the extractor modules and the edgeKeys deduplication in GraphBuilder—the system can detect meaningful code changes instantly while preserving expensive LLM annotations when source remains unchanged.
Summary
- Tree‑Sitter provides deterministic AST parsing via
TreeSitterPlugin, ensuring repeatable structural analysis. - Pure extractors in
builtinExtractorswalk the AST without side effects, generating consistentStructuralAnalysisobjects. - Stable identifiers like
function:src/app.ts:myFuncand theedgeKeysdeduplication set inGraphBuilderguarantee identical graph topology for identical source. - LLM enrichment occurs only after structural nodes exist, adding
summary,tags, andcomplexitywithout altering node IDs or edges. - Separation of concerns allows the structural graph to be cached while semantic data refreshes independently.
Frequently Asked Questions
How does Understand Anything ensure the graph structure is deterministic?
The tool uses Tree‑Sitter's grammar‑driven parser to produce identical ASTs for identical source code. The GraphBuilder class generates stable node IDs based on file paths and symbol names (e.g., function:src/app.ts:myFunc) and deduplicates edges using an edgeKeys Set. These mechanisms ensure that parsing the same file twice yields the exact same node and edge topology.
What prevents LLM randomness from affecting the graph structure?
Semantic enrichment occurs only after the structural graph is constructed. The LLM analyzer in llm-analyzer.ts injects metadata like summaries and tags into existing nodes but cannot create new nodes or modify edge relationships. Because the LLM output is parsed and stored as node properties—not as structural elements—the graph topology remains stable regardless of LLM temperature or response variations.
Which files handle the language‑specific AST extraction?
Language detection occurs in language-registry.ts, while the actual parsing and extraction logic resides in tree-sitter-plugin.ts and the extractor modules referenced in extractors/index.ts. Each extractor is a pure function that walks the Tree‑sitter AST and returns a StructuralAnalysis object containing functions, classes, imports, and exports.
Can the deterministic graph be used without LLM enrichment?
Yes. The GraphBuilder can construct a complete knowledge graph using only the structural analysis from TreeSitterPlugin.analyzeFile(). The LLM step is optional and only adds human‑readable intent. This makes the tool suitable for environments where deterministic, reproducible code analysis is required without AI‑generated content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →