# Tree-Sitter + LLM Hybrid Architecture: How Understand-Anything Extracts Structural and Semantic Code

> Discover the tree-sitter + LLM hybrid architecture. This powerful combination extracts precise structural and semantic code information, creating a unified knowledge graph for better code understanding.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: architecture
- Published: 2026-05-31

---

**The tree-sitter + LLM hybrid architecture combines deterministic AST parsing via Tree-Sitter with generative reasoning through a large language model to produce a unified knowledge graph containing both precise code relationships and human-readable semantic annotations.**

The **Understand-Anything** engine employs a sophisticated tree-sitter + LLM hybrid architecture to analyze source code at two complementary levels. This dual-layer approach first extracts deterministic structural facts through Tree-Sitter's language-specific grammars, then enriches that skeleton with semantic context generated by an LLM. The result is a comprehensive code understanding system that captures both *what* the code contains and *what it means*.

## Structural Extraction with Tree-Sitter

The structural layer relies on `TreeSitterPlugin` located at [`understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts). This plugin performs fast, deterministic parsing by loading Tree-Sitter grammars compiled to WebAssembly, enabling language-aware analysis without LLM latency.

### Grammar Initialization and WASM Loading

During initialization, `TreeSitterPlugin.init()` bootstraps the WebAssembly runtime and loads the appropriate grammar for each configured language (e.g., `tree-sitter-typescript`, `tree-sitter-python`). This occurs at lines 24‑38 and 124‑140 of the plugin file.

```typescript
const mod = await import("web-tree-sitter");
const ParserCls = mod.Parser;
await ParserCls.init();                 // ← bootstraps WASM
const lang = await LanguageCls.load(wasmPath);
this._languages.set(config.id, lang);

```

### Parser Selection and AST Traversal

When analyzing a file, the plugin maps file extensions to language keys and instantiates a parser configured with the pre-loaded grammar. Language-specific extractors—such as [`typescript-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/typescript-extractor.ts) or [`python-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/python-extractor.ts)—then walk the AST to produce a `StructuralAnalysis` object.

```typescript
const ext = extname(filePath).toLowerCase();
const langKey = this._extensionToLang.get(ext) ?? null;
const parser = new this._ParserClass();
parser.setLanguage(this._languages.get(langKey)!);
const analysis = extractor.extractStructure(tree.rootNode);

```

The resulting `StructuralAnalysis` contains deterministic data: function names and line ranges, class hierarchies, import/export statements, and resolved caller-callee pairs for call graph construction.

## Semantic Extraction with an LLM

While Tree-Sitter provides the skeleton, the semantic layer adds meaning. Implemented in [`understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts), this component uses prompt-driven reasoning to generate human-readable summaries and complexity assessments.

### Prompt Construction

The `buildFileAnalysisPrompt` function (lines 19‑42) assembles a structured instruction that includes the file path, full source content, and a checklist of required JSON fields.

```typescript
export function buildFileAnalysisPrompt(filePath, content, projectContext) {
  return `You are a code analysis assistant…
File: ${filePath}
\`\`\`
${content}
\`\`\`
…Return a JSON object…`;
}

```

### Response Parsing and Typed Output

After receiving the LLM response, the `extractJson` helper strips markdown fences and parses the result. The output conforms to the `LLMFileAnalysis` interface, ensuring type safety for downstream consumers.

```typescript
const jsonStr = extractJson(response);
const parsed = JSON.parse(jsonStr);

```

The `LLMFileAnalysis` interface defines the semantic schema:

```typescript
export interface LLMFileAnalysis {
  fileSummary: string;
  tags: string[];
  complexity: "simple" | "moderate" | "complex";
  functionSummaries: Record<string, string>;
  classSummaries: Record<string, string>;
  languageNotes?: string;
}

```

## Merging Structural and Semantic Data in the Knowledge Graph

The `GraphBuilder` class, defined in [`understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts), serves as the fusion point for the tree-sitter + LLM hybrid architecture. It consumes the deterministic `StructuralAnalysis` from Tree-Sitter and the semantic metadata from the LLM to construct a unified knowledge graph.

### The GraphBuilder Integration

The `addFileWithAnalysis` method stitches both data streams together. File nodes store LLM-generated summaries, tags, and complexity ratings, while function and class nodes receive per-entity semantic descriptions mapped by name from the structural analysis.

```typescript
builder.addFileWithAnalysis(filePath, structuralAnalysis, {
  summary: llmMeta.fileSummary,
  tags: llmMeta.tags,
  complexity: llmMeta.complexity,
  fileSummary: llmMeta.fileSummary,
  summaries: { ...llmMeta.functionSummaries, ...llmMeta.classSummaries },
});

```

### Constructing Relationships

Import edges (`addImportEdge`) and call edges (`addCallEdge`) are derived exclusively from the structural analysis, providing a precise, language-accurate map of code relationships. These edges connect nodes that have been enriched with LLM-generated summaries, creating a graph that supports both navigation by structure and comprehension by semantics.

## Complete Implementation Example

The following end-to-end example demonstrates how to initialize the architecture, extract both analysis types, and build the unified graph:

```typescript
import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";
import { GraphBuilder } from "./analyzer/graph-builder.js";
import { buildFileAnalysisPrompt, parseFileAnalysisResponse } from "./analyzer/llm-analyzer.js";

// 1️⃣ Initialise the structural parser (load TS + JS grammars)
const treeSitter = new TreeSitterPlugin();
await treeSitter.init();

// 2️⃣ Read a source file
const filePath = "src/utils.ts";
const content = await Deno.readTextFile(filePath);

// 3️⃣ Structural analysis (functions, classes, imports)
const structural = treeSitter.analyzeFile(filePath, content);

// 4️⃣ Generate an LLM prompt and obtain a response
const prompt = buildFileAnalysisPrompt(filePath, content, "Utility library for string helpers");
const llmResponse = await callYourLLM(prompt);       // <-- your LLM provider
const semantic = parseFileAnalysisResponse(llmResponse);

// 5️⃣ Build the knowledge graph
const builder = new GraphBuilder("my‑project", "deadbeef");
builder.addFileWithAnalysis(filePath, structural, {
  summary: semantic?.fileSummary ?? "",
  tags: semantic?.tags ?? [],
  complexity: semantic?.complexity ?? "moderate",
  fileSummary: semantic?.fileSummary ?? "",
  summaries: {
    ...semantic?.functionSummaries,
    ...semantic?.classSummaries,
  },
});
const graph = builder.build();

```

## Key Source Files

The tree-sitter + LLM hybrid architecture is implemented across the following core modules:

- **[`understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts)** – Core Tree-Sitter driver that loads WASM grammars and orchestrates parsing.
- **[`understand-anything-plugin/packages/core/src/plugins/extractors/typescript-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/extractors/typescript-extractor.ts)** – Language-specific extractor demonstrating AST walking for TypeScript.
- **[`understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts)** – Prompt builders, response parsers, and the `LLMFileAnalysis` interface.
- **[`understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts)** – Graph construction logic that merges structural and semantic data streams.
- **[`understand-anything-plugin/packages/core/src/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/index.ts)** – Public façade re-exporting LLM types for downstream consumers.

## Summary

- The **tree-sitter + LLM hybrid architecture** separates deterministic structure extraction from generative semantic analysis.
- **Tree-Sitter** provides precise AST parsing via WebAssembly grammars, yielding function definitions, class hierarchies, and call graphs.
- The **LLM layer** generates human-readable summaries, complexity scores, and tags through prompt-driven reasoning.
- **GraphBuilder** fuses both streams into a unified knowledge graph, enabling tools to navigate code relationships while understanding their purpose.
- All components are modular, allowing the structural analyzer to operate independently when LLM latency or cost is a concern.

## Frequently Asked Questions

### Why combine Tree-Sitter with an LLM instead of using just one approach?

Tree-Sitter delivers deterministic, fast, and language-accurate structural data such as precise line ranges and call graphs, which LLMs often hallucinate or approximate. Conversely, LLMs excel at generating human-readable summaries and semantic context that static analysis cannot infer. The hybrid architecture leverages the strengths of each: Tree-Sitter provides the reliable skeleton, while the LLM adds the interpretive layer.

### How does the GraphBuilder handle conflicts between structural and semantic data?

The architecture is designed to avoid conflicts by assigning distinct responsibilities. `GraphBuilder` uses structural data exclusively for relationship edges (imports, calls) and node positioning, while LLM data populates metadata fields (summaries, tags). Since these domains do not overlap, the merge process in `addFileWithAnalysis` is strictly additive, enriching structural nodes with semantic attributes without altering the underlying graph topology.

### What programming languages does the Tree-Sitter plugin support?

The plugin supports any language with a Tree-Sitter grammar compiled to WebAssembly. The repository specifically includes extractors for **TypeScript** and **Python**, but the modular design in [`tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/tree-sitter-plugin.ts) allows extending support by loading additional WASM grammars (e.g., `tree-sitter-rust`, `tree-sitter-go`) and registering corresponding file extension mappings.

### Is the LLM analysis performed locally or via an external API?

The [`llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/llm-analyzer.ts) module is provider-agnostic. It constructs prompts and parses responses, but delegates the actual LLM call to the consumer's implementation (illustrated by the `callYourLLM` placeholder in the examples). This design supports both local models (via Ollama or similar) and remote APIs (OpenAI, Anthropic, etc.) without modifying the core architecture.