# How Tree-Sitter and LLM Hybrid Analysis Works in Understand-Anything

> Discover how Understand-Anything leverages tree-sitter and LLM hybrid analysis to create a navigable knowledge graph of your codebase. Explore code with unprecedented depth.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-02

---

**Understand-Anything combines precise AST parsing via tree-sitter with generative LLM insights to build a rich, navigable knowledge graph of any codebase.**

This hybrid architecture pairs deterministic structural extraction with semantic understanding, enabling the tool to map exact code relationships while generating human-readable summaries. The approach leverages `TreeSitterPlugin` for fast, language-aware parsing and [`llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/llm-analyzer.ts) for contextual metadata generation, merging both streams into a unified graph structure as implemented in Lum1104/Understand-Anything.

## Tree-Sitter Structural Analysis Layer

The foundation of the hybrid system relies on **tree-sitter** to produce exact, language-specific AST representations. This layer extracts functions, classes, imports, and exports with deterministic precision, independent of the LLM's interpretation.

### Loading WASM Grammars and Parser Initialization

`TreeSitterPlugin` initializes by reading the `treeSitter` field from each `LanguageConfig` and loading the associated `.wasm` grammar file via `web-tree-sitter`. For a given source file, the plugin selects the appropriate language parser based on file extension (`.ts`, `.js`, `.py`, etc.) and creates a dedicated `Parser` instance.

```typescript
import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";

const plugin = new TreeSitterPlugin();
await plugin.init(); // Loads WASM grammars into memory

const analysis = plugin.analyzeFile(
  "src/example.ts",
  readFileSync("src/example.ts", "utf-8")
);
console.log(analysis.functions); // → [{ name: "foo", … }]

```

*Source*: [[`understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts)

### AST Extraction and Language-Specific Walkers

Once the AST is built, language-specific extractors (located in [`extractors/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/extractors/index.ts)) walk the tree and return a `StructuralAnalysis` object. These extractors handle TypeScript, JavaScript, Python, Go, and other supported languages, normalizing the AST into a consistent schema containing functions, classes, and dependency information.

## LLM-Driven Semantic Analysis Layer

While tree-sitter captures *what* exists in the code, the **LLM layer** explains *what it means*. This component generates high-level summaries, complexity ratings, and semantic tags that pure static analysis cannot deduce.

### JSON-Only Prompt Construction

`buildFileAnalysisPrompt` (in [`llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/llm-analyzer.ts)) constructs strict instructions that embed raw file content and optional project context. The prompt explicitly requires the LLM to return **only** a JSON object containing fields like `fileSummary`, `tags`, and `complexity`.

```typescript
import {
  buildFileAnalysisPrompt,
  parseFileAnalysisResponse,
} from "./analyzer/llm-analyzer.js";

const prompt = buildFileAnalysisPrompt(
  "src/example.ts",
  fileContent,
  "A small CLI utility"
);
// → Prompt sent to LLM via external skill

```

### Response Parsing and Type Validation

`parseFileAnalysisResponse` extracts the JSON block—even if wrapped in markdown fences—and validates required fields including `fileSummary`, `tags`, `complexity`, `functionSummaries`, and `classSummaries`. This ensures type safety before the metadata enters the knowledge graph.

```typescript
const analysis = parseFileAnalysisResponse(rawResponse);
console.log(analysis?.tags);   // e.g., ["utility", "async"]
console.log(analysis?.complexity); // "low" | "moderate" | "high"

```

*Source*: [[`understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/llm-analyzer.ts)

## Merging Structural and Semantic Data in the Graph Builder

`GraphBuilder` (in [`graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-builder.ts)) serves as the integration point, receiving both the `StructuralAnalysis` from tree-sitter and the `FileAnalysisMeta` from the LLM. The builder:

- Creates **file nodes** containing LLM-derived summaries, tags, and complexity ratings
- Adds **function** and **class** child nodes, attaching LLM-generated summaries when available
- Wires **contains**, **imports**, and **calls** edges based strictly on structural data

```typescript
const structural = treeSitterPlugin.analyzeFile(path, content);
const llmMeta = await getLLMFileAnalysis(path, content);

graphBuilder.addFileWithAnalysis(
  path,
  structural,
  {
    fileSummary: llmMeta.fileSummary,
    tags: llmMeta.tags,
    complexity: llmMeta.complexity,
    summaries: {
      ...llmMeta.functionSummaries,
      ...llmMeta.classSummaries,
    },
  }
);

```

*Source*: [[`understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts)

## Change Detection via Structural Fingerprints

The system uses tree-sitter for intelligent **change detection** through [`fingerprint.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/fingerprint.ts). When re-analyzing a project, the tool generates a *structural fingerprint* based on function signatures, class definitions, and import statements derived from the AST.

If the fingerprint remains unchanged, the system records a **cosmetic** change; structural differences trigger a full re-analysis. Files lacking tree-sitter support fall back to content-hash fingerprints, conservatively flagging all modifications as structural.

*Source*: [[`understand-anything-plugin/packages/core/src/fingerprint.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/fingerprint.ts)](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/fingerprint.ts)

## Practical Implementation Examples

### Parsing Source Files with TreeSitterPlugin

```typescript
import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";
import { readFileSync } from "node:fs";

const plugin = new TreeSitterPlugin();
await plugin.init();

const pyPath = "scripts/example.py";
const pyContent = readFileSync(pyPath, "utf-8");
const analysis = plugin.analyzeFile(pyPath, pyContent);

console.log("Functions:", analysis.functions.map(f => f.name));
console.log("Imports:", analysis.imports.map(i => i.source));

```

### Generating Metadata with llm-analyzer.ts

```typescript
import {
  buildFileAnalysisPrompt,
  parseFileAnalysisResponse,
} from "./analyzer/llm-analyzer.js";
import { sendToLLM } from "./llm-client.js";

async function llmFileMeta(filePath: string, content: string) {
  const prompt = buildFileAnalysisPrompt(filePath, content, "Generic web app");
  const raw = await sendToLLM(prompt);
  return parseFileAnalysisResponse(raw);
}

```

### Constructing the Knowledge Graph

```typescript
import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";
import { GraphBuilder } from "./analyzer/graph-builder.js";
import { llmFileMeta } from "./llm-wrapper.js";

async function buildProjectGraph(projectName, gitHash, files) {
  const treePlugin = new TreeSitterPlugin();
  await treePlugin.init();

  const builder = new GraphBuilder(projectName, gitHash);

  for (const { path, content } of files) {
    const structural = treePlugin.analyzeFile(path, content);
    const llmMeta = await llmFileMeta(path, content);

    builder.addFileWithAnalysis(
      path,
      structural,
      {
        fileSummary: llmMeta?.fileSummary ?? "",
        tags: llmMeta?.tags ?? [],
        complexity: llmMeta?.complexity ?? "moderate",
        summaries: {
          ...llmMeta?.functionSummaries,
          ...llmMeta?.classSummaries,
        },
      }
    );
  }

  return builder.build(); // → KnowledgeGraph ready for dashboard
}

```

## Summary

- **Tree-sitter** provides deterministic, language-aware AST parsing via `TreeSitterPlugin`, extracting exact structural relationships (functions, classes, imports) from source files.
- **LLM analysis** generates semantic metadata—summaries, tags, and complexity ratings—through strictly formatted JSON prompts handled in [`llm-analyzer.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/llm-analyzer.ts).
- **GraphBuilder** merges both data streams into a unified knowledge graph, combining precise code topology with human-readable context.
- **Fingerprinting** leverages tree-sitter output to distinguish cosmetic changes from structural modifications, optimizing re-analysis performance.

## Frequently Asked Questions

### How does Understand-Anything handle multiple programming languages?

`TreeSitterPlugin` uses a `LanguageRegistry` to map file extensions to language-specific WASM grammars. Each language has a dedicated extractor in [`extractors/index.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/extractors/index.ts) that normalizes the AST into a common `StructuralAnalysis` format, allowing the system to treat TypeScript, Python, Go, and other languages uniformly in the graph construction phase.

### Why use both tree-sitter and LLM instead of just one approach?

**Tree-sitter** delivers millisecond-level parsing speed and deterministic accuracy for code structure, while **LLMs** provide contextual understanding that static analysis cannot capture (e.g., "this function handles authentication"). Using both ensures the knowledge graph contains exact dependency relationships *and* semantic meaning necessary for features like automated code tours and natural language querying.

### What happens if the LLM returns malformed JSON?

`parseFileAnalysisResponse` includes robust extraction logic that detects JSON blocks even when wrapped in markdown code fences. If validation fails or required fields are missing, the system falls back to default values (empty summaries, "moderate" complexity) and logs the error, ensuring the graph build process continues without interrupting the analysis pipeline.

### How does the system detect which files changed between analyses?

[`fingerprint.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/fingerprint.ts) generates structural fingerprints by hashing the tree-sitter output (function signatures, class names, imports). This allows the system to distinguish between meaningful structural changes (requiring re-analysis) and cosmetic formatting changes (preserving cached metadata), significantly reducing redundant LLM API calls on large codebases.