Tree-Sitter + LLM Hybrid Architecture: How Understand-Anything Extracts Structural and Semantic Code
The tree-sitter + LLM hybrid architecture combines deterministic AST parsing via Tree-Sitter with generative reasoning through a large language model to produce a unified knowledge graph containing both precise code relationships and human-readable semantic annotations.
The Understand-Anything engine employs a sophisticated tree-sitter + LLM hybrid architecture to analyze source code at two complementary levels. This dual-layer approach first extracts deterministic structural facts through Tree-Sitter's language-specific grammars, then enriches that skeleton with semantic context generated by an LLM. The result is a comprehensive code understanding system that captures both what the code contains and what it means.
Structural Extraction with Tree-Sitter
The structural layer relies on TreeSitterPlugin located at understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts. This plugin performs fast, deterministic parsing by loading Tree-Sitter grammars compiled to WebAssembly, enabling language-aware analysis without LLM latency.
Grammar Initialization and WASM Loading
During initialization, TreeSitterPlugin.init() bootstraps the WebAssembly runtime and loads the appropriate grammar for each configured language (e.g., tree-sitter-typescript, tree-sitter-python). This occurs at lines 24‑38 and 124‑140 of the plugin file.
const mod = await import("web-tree-sitter");
const ParserCls = mod.Parser;
await ParserCls.init(); // ← bootstraps WASM
const lang = await LanguageCls.load(wasmPath);
this._languages.set(config.id, lang);
Parser Selection and AST Traversal
When analyzing a file, the plugin maps file extensions to language keys and instantiates a parser configured with the pre-loaded grammar. Language-specific extractors—such as typescript-extractor.ts or python-extractor.ts—then walk the AST to produce a StructuralAnalysis object.
const ext = extname(filePath).toLowerCase();
const langKey = this._extensionToLang.get(ext) ?? null;
const parser = new this._ParserClass();
parser.setLanguage(this._languages.get(langKey)!);
const analysis = extractor.extractStructure(tree.rootNode);
The resulting StructuralAnalysis contains deterministic data: function names and line ranges, class hierarchies, import/export statements, and resolved caller-callee pairs for call graph construction.
Semantic Extraction with an LLM
While Tree-Sitter provides the skeleton, the semantic layer adds meaning. Implemented in understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts, this component uses prompt-driven reasoning to generate human-readable summaries and complexity assessments.
Prompt Construction
The buildFileAnalysisPrompt function (lines 19‑42) assembles a structured instruction that includes the file path, full source content, and a checklist of required JSON fields.
export function buildFileAnalysisPrompt(filePath, content, projectContext) {
return `You are a code analysis assistant…
File: ${filePath}
\`\`\`
${content}
\`\`\`
…Return a JSON object…`;
}
Response Parsing and Typed Output
After receiving the LLM response, the extractJson helper strips markdown fences and parses the result. The output conforms to the LLMFileAnalysis interface, ensuring type safety for downstream consumers.
const jsonStr = extractJson(response);
const parsed = JSON.parse(jsonStr);
The LLMFileAnalysis interface defines the semantic schema:
export interface LLMFileAnalysis {
fileSummary: string;
tags: string[];
complexity: "simple" | "moderate" | "complex";
functionSummaries: Record<string, string>;
classSummaries: Record<string, string>;
languageNotes?: string;
}
Merging Structural and Semantic Data in the Knowledge Graph
The GraphBuilder class, defined in understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts, serves as the fusion point for the tree-sitter + LLM hybrid architecture. It consumes the deterministic StructuralAnalysis from Tree-Sitter and the semantic metadata from the LLM to construct a unified knowledge graph.
The GraphBuilder Integration
The addFileWithAnalysis method stitches both data streams together. File nodes store LLM-generated summaries, tags, and complexity ratings, while function and class nodes receive per-entity semantic descriptions mapped by name from the structural analysis.
builder.addFileWithAnalysis(filePath, structuralAnalysis, {
summary: llmMeta.fileSummary,
tags: llmMeta.tags,
complexity: llmMeta.complexity,
fileSummary: llmMeta.fileSummary,
summaries: { ...llmMeta.functionSummaries, ...llmMeta.classSummaries },
});
Constructing Relationships
Import edges (addImportEdge) and call edges (addCallEdge) are derived exclusively from the structural analysis, providing a precise, language-accurate map of code relationships. These edges connect nodes that have been enriched with LLM-generated summaries, creating a graph that supports both navigation by structure and comprehension by semantics.
Complete Implementation Example
The following end-to-end example demonstrates how to initialize the architecture, extract both analysis types, and build the unified graph:
import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";
import { GraphBuilder } from "./analyzer/graph-builder.js";
import { buildFileAnalysisPrompt, parseFileAnalysisResponse } from "./analyzer/llm-analyzer.js";
// 1️⃣ Initialise the structural parser (load TS + JS grammars)
const treeSitter = new TreeSitterPlugin();
await treeSitter.init();
// 2️⃣ Read a source file
const filePath = "src/utils.ts";
const content = await Deno.readTextFile(filePath);
// 3️⃣ Structural analysis (functions, classes, imports)
const structural = treeSitter.analyzeFile(filePath, content);
// 4️⃣ Generate an LLM prompt and obtain a response
const prompt = buildFileAnalysisPrompt(filePath, content, "Utility library for string helpers");
const llmResponse = await callYourLLM(prompt); // <-- your LLM provider
const semantic = parseFileAnalysisResponse(llmResponse);
// 5️⃣ Build the knowledge graph
const builder = new GraphBuilder("my‑project", "deadbeef");
builder.addFileWithAnalysis(filePath, structural, {
summary: semantic?.fileSummary ?? "",
tags: semantic?.tags ?? [],
complexity: semantic?.complexity ?? "moderate",
fileSummary: semantic?.fileSummary ?? "",
summaries: {
...semantic?.functionSummaries,
...semantic?.classSummaries,
},
});
const graph = builder.build();
Key Source Files
The tree-sitter + LLM hybrid architecture is implemented across the following core modules:
understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts– Core Tree-Sitter driver that loads WASM grammars and orchestrates parsing.understand-anything-plugin/packages/core/src/plugins/extractors/typescript-extractor.ts– Language-specific extractor demonstrating AST walking for TypeScript.understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts– Prompt builders, response parsers, and theLLMFileAnalysisinterface.understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts– Graph construction logic that merges structural and semantic data streams.understand-anything-plugin/packages/core/src/index.ts– Public façade re-exporting LLM types for downstream consumers.
Summary
- The tree-sitter + LLM hybrid architecture separates deterministic structure extraction from generative semantic analysis.
- Tree-Sitter provides precise AST parsing via WebAssembly grammars, yielding function definitions, class hierarchies, and call graphs.
- The LLM layer generates human-readable summaries, complexity scores, and tags through prompt-driven reasoning.
- GraphBuilder fuses both streams into a unified knowledge graph, enabling tools to navigate code relationships while understanding their purpose.
- All components are modular, allowing the structural analyzer to operate independently when LLM latency or cost is a concern.
Frequently Asked Questions
Why combine Tree-Sitter with an LLM instead of using just one approach?
Tree-Sitter delivers deterministic, fast, and language-accurate structural data such as precise line ranges and call graphs, which LLMs often hallucinate or approximate. Conversely, LLMs excel at generating human-readable summaries and semantic context that static analysis cannot infer. The hybrid architecture leverages the strengths of each: Tree-Sitter provides the reliable skeleton, while the LLM adds the interpretive layer.
How does the GraphBuilder handle conflicts between structural and semantic data?
The architecture is designed to avoid conflicts by assigning distinct responsibilities. GraphBuilder uses structural data exclusively for relationship edges (imports, calls) and node positioning, while LLM data populates metadata fields (summaries, tags). Since these domains do not overlap, the merge process in addFileWithAnalysis is strictly additive, enriching structural nodes with semantic attributes without altering the underlying graph topology.
What programming languages does the Tree-Sitter plugin support?
The plugin supports any language with a Tree-Sitter grammar compiled to WebAssembly. The repository specifically includes extractors for TypeScript and Python, but the modular design in tree-sitter-plugin.ts allows extending support by loading additional WASM grammars (e.g., tree-sitter-rust, tree-sitter-go) and registering corresponding file extension mappings.
Is the LLM analysis performed locally or via an external API?
The llm-analyzer.ts module is provider-agnostic. It constructs prompts and parses responses, but delegates the actual LLM call to the consumer's implementation (illustrated by the callYourLLM placeholder in the examples). This design supports both local models (via Ollama or similar) and remote APIs (OpenAI, Anthropic, etc.) without modifying the core architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →