# How Egonex AI Combines Tree-sitter and LLMs for Deep Code Analysis

> Discover how Egonex AI uses Tree-sitter and LLMs to analyze code deeply. Learn about their three-stage pipeline for automated code understanding across many languages.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-17

---

**Egonex AI merges Tree-sitter's fast incremental parsing with LLM semantic reasoning through a three-stage pipeline that transforms concrete syntax trees into an enriched knowledge graph, enabling automated code understanding across dozens of languages.**

The *Understand Anything* platform by Egonex AI addresses large-scale code comprehension by uniting deterministic syntactic analysis with generative AI. This architecture leverages Tree-sitter's WebAssembly parser to extract precise code structures, then feeds those structures to an LLM for semantic enrichment, producing an interactive knowledge graph that powers the dashboard's code tours and security insights.

## The Three-Stage Pipeline Architecture

### Stage 1: Parsing with Tree-sitter

The pipeline begins in [`packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/tree-sitter-plugin.ts), where the **Web-tree-sitter** library loads language-specific grammars registered in [`packages/core/src/languages/language-registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/languages/language-registry.ts) and managed by [`packages/core/src/registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/registry.ts). For each source file, Tree-sitter generates a **Concrete Syntax Tree (CST)** that captures exact syntactic relationships, including scopes, imports, and identifiers. This WebAssembly-based parser handles incremental updates efficiently, re-parsing only changed fragments when files are modified.

### Stage 2: Graph Construction

The raw CST flows into [`packages/core/src/analyzer/graph-builder.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/analyzer/graph-builder.ts), which traverses the tree to extract nodes, edges, and metadata. This module normalizes language-specific ASTs into a unified schema defined in [`packages/core/src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/types.ts), creating a language-agnostic knowledge graph that represents the entire repository structure. The graph captures cross-file relationships and code boundaries that pure text analysis would miss.

### Stage 3: LLM Enrichment

Finally, [`packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/analyzer/llm-analyzer.ts) serializes graph fragments into prompts for the **LLM Analyzer**. Using few-shot prompting techniques, the model infers high-level semantic concepts—such as identifying React components or security-critical entry points—and generates human-readable summaries. The LLM's output is stitched back into the graph, augmenting each node with `tags`, `summary`, and `confidence` scores before validation by [`packages/core/src/schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/schema.ts).

## Why Tree-sitter and LLMs Complement Each Other

This hybrid approach leverages the strengths of both technologies while mitigating their individual limitations:

- **Tree-sitter for syntactic fidelity**: Guarantees exact, language-specific ASTs and CSTs for supported languages including JavaScript, Python, Java, and Kotlin. The parser's incremental capabilities enable rapid updates when files change.

- **LLM for semantic reasoning**: Supplies intent, design patterns, and cross-file relationships that pure parsing cannot detect. The model interprets the generic CST format through language-agnostic prompt templates.

- **Speed and efficiency**: Tree-sitter re-parses only modified fragments, while the LLM receives batched requests containing only changed sub-trees, minimizing token usage and API costs.

- **Extensible architecture**: The plugin registry in [`packages/core/src/registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/registry.ts) isolates language-specific grammars, while new analysis capabilities are added by extending the prompt library in the analyzer modules.

## End-to-End Data Flow

The complete analysis pipeline follows five distinct phases orchestrated by the core engine:

1. **Discovery**: [`plugins/discovery.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/plugins/discovery.ts) scans the project root to identify supported source files based on registered language extensions.

2. **Parsing**: [`tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/tree-sitter-plugin.ts) invokes the appropriate Tree-sitter grammar for each file, emitting standardized AST representations.

3. **Normalization**: [`normalize-graph.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/normalize-graph.ts) transforms raw ASTs into the unified graph schema, ensuring consistent node types across languages.

4. **LLM Enrichment**: [`llm-analyzer.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/llm-analyzer.ts) formats graph nodes into prompts, calls the configured LLM (defaulting to GPT-4o-mini), and injects semantic metadata back into the graph.

5. **Persistence**: The enriched graph is validated by [`schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/schema.ts) and persisted via [`persistence/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/persistence/index.ts), then consumed by [`packages/dashboard/src/store.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/dashboard/src/store.ts) for UI rendering.

The dashboard specifically imports only **browser-safe** sub-path exports (`./search`, `./types`, `./schema`) to prevent the heavyweight Tree-sitter WASM bundle from executing in the browser, keeping the UI lightweight and responsive.

## Programmatic Usage Example

You can trigger the full pipeline programmatically using the `analyzeProject` function from the core package. This example requires Node.js >= 22 and assumes a PNPM workspace setup:

```typescript
import { analyzeProject } from '@understand-anything/core';
import { writeFile } from 'fs/promises';

(async () => {
  // Point to the repository root (absolute path)
  const repoRoot = '/path/to/your/project';

  // Run the complete analysis pipeline:
  // discovery → tree-sitter parsing → graph building → LLM enrichment
  const graph = await analyzeProject(repoRoot, {
    llm: { model: 'gpt-4o-mini', temperature: 0.0 },
    languages: ['javascript', 'typescript', 'java'],
  });

  // Access enriched nodes with semantic metadata
  console.log('First node summary:', graph.nodes[0]?.summary);

  // Persist for dashboard consumption
  await writeFile(
    `${repoRoot}/.understand-anything/knowledge-graph.json`,
    JSON.stringify(graph, null, 2)
  );
})();

```

After building the core package and running the dashboard:

```bash
pnpm --filter @understand-anything/core build
pnpm --filter @understand-anything/dashboard dev

```

The dashboard will read the generated [`knowledge-graph.json`](https://github.com/Egonex-AI/Understand-Anything/blob/main/knowledge-graph.json), validate it against [`schema.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/schema.ts), and render an interactive graph where each node displays the LLM-generated summary and semantic tags.

## Summary

- **Tree-sitter integration** in [`tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/tree-sitter-plugin.ts) provides fast, incremental parsing via WebAssembly, generating precise CSTs for multiple languages.
- **Graph construction** in [`graph-builder.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/graph-builder.ts) normalizes these trees into a unified, language-agnostic knowledge graph representing the full codebase.
- **LLM enrichment** in [`llm-analyzer.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/llm-analyzer.ts) adds semantic understanding through few-shot prompting, generating tags, summaries, and security classifications.
- **Incremental updates** ensure only changed fragments are re-parsed and re-analyzed, optimizing both compute and token costs.
- **Browser-safe architecture** keeps the heavy Tree-sitter WASM in the Node.js backend while serving lightweight data to the React dashboard.

## Frequently Asked Questions

### What role does Tree-sitter play in the Egonex AI pipeline?

Tree-sitter serves as the deterministic syntactic foundation. Implemented in [`packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/tree-sitter-plugin.ts) using the Web-tree-sitter library, it parses source files into Concrete Syntax Trees that capture exact grammatical relationships, scopes, and imports. This guarantees syntactic fidelity that LLMs alone cannot provide, while its incremental parsing capability enables rapid updates when code changes.

### How does the LLM analyzer enrich the knowledge graph?

The LLM analyzer in [`packages/core/src/analyzer/llm-analyzer.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/analyzer/llm-analyzer.ts) receives serialized graph fragments and uses few-shot prompts to infer high-level semantic concepts. It identifies design patterns, security-critical entry points, and component types, then augments graph nodes with `tags`, human-readable `summaries`, and `confidence` scores. This semantic layer transforms raw syntax into actionable insights like automated code tours and vulnerability detection.

### Can the system handle incremental code updates efficiently?

Yes. Tree-sitter's incremental parser re-parses only changed file fragments rather than entire files, while the LLM analyzer receives batched prompts containing only modified sub-trees. This architecture minimizes both computational overhead and LLM token usage, enabling real-time updates to the knowledge graph as developers write code.

### Is the dashboard browser-compatible with Tree-sitter?

The dashboard intentionally avoids running Tree-sitter in the browser. Instead, it imports only browser-safe sub-paths (`./search`, `./types`, `./schema`) and consumes pre-computed knowledge graphs persisted by the backend. The heavyweight Tree-sitter WASM bundle executes exclusively in the Node.js backend during the analysis phase, ensuring the React UI remains lightweight and responsive.