How Egonex AI Combines Tree-sitter and LLMs for Deep Code Analysis

Egonex AI merges Tree-sitter's fast incremental parsing with LLM semantic reasoning through a three-stage pipeline that transforms concrete syntax trees into an enriched knowledge graph, enabling automated code understanding across dozens of languages.

The Understand Anything platform by Egonex AI addresses large-scale code comprehension by uniting deterministic syntactic analysis with generative AI. This architecture leverages Tree-sitter's WebAssembly parser to extract precise code structures, then feeds those structures to an LLM for semantic enrichment, producing an interactive knowledge graph that powers the dashboard's code tours and security insights.

The Three-Stage Pipeline Architecture

Stage 1: Parsing with Tree-sitter

The pipeline begins in packages/core/src/plugins/tree-sitter-plugin.ts, where the Web-tree-sitter library loads language-specific grammars registered in packages/core/src/languages/language-registry.ts and managed by packages/core/src/registry.ts. For each source file, Tree-sitter generates a Concrete Syntax Tree (CST) that captures exact syntactic relationships, including scopes, imports, and identifiers. This WebAssembly-based parser handles incremental updates efficiently, re-parsing only changed fragments when files are modified.

Stage 2: Graph Construction

The raw CST flows into packages/core/src/analyzer/graph-builder.ts, which traverses the tree to extract nodes, edges, and metadata. This module normalizes language-specific ASTs into a unified schema defined in packages/core/src/types.ts, creating a language-agnostic knowledge graph that represents the entire repository structure. The graph captures cross-file relationships and code boundaries that pure text analysis would miss.

Stage 3: LLM Enrichment

Finally, packages/core/src/analyzer/llm-analyzer.ts serializes graph fragments into prompts for the LLM Analyzer. Using few-shot prompting techniques, the model infers high-level semantic concepts—such as identifying React components or security-critical entry points—and generates human-readable summaries. The LLM's output is stitched back into the graph, augmenting each node with tags, summary, and confidence scores before validation by packages/core/src/schema.ts.

Why Tree-sitter and LLMs Complement Each Other

This hybrid approach leverages the strengths of both technologies while mitigating their individual limitations:

  • Tree-sitter for syntactic fidelity: Guarantees exact, language-specific ASTs and CSTs for supported languages including JavaScript, Python, Java, and Kotlin. The parser's incremental capabilities enable rapid updates when files change.

  • LLM for semantic reasoning: Supplies intent, design patterns, and cross-file relationships that pure parsing cannot detect. The model interprets the generic CST format through language-agnostic prompt templates.

  • Speed and efficiency: Tree-sitter re-parses only modified fragments, while the LLM receives batched requests containing only changed sub-trees, minimizing token usage and API costs.

  • Extensible architecture: The plugin registry in packages/core/src/registry.ts isolates language-specific grammars, while new analysis capabilities are added by extending the prompt library in the analyzer modules.

End-to-End Data Flow

The complete analysis pipeline follows five distinct phases orchestrated by the core engine:

  1. Discovery: plugins/discovery.ts scans the project root to identify supported source files based on registered language extensions.

  2. Parsing: tree-sitter-plugin.ts invokes the appropriate Tree-sitter grammar for each file, emitting standardized AST representations.

  3. Normalization: normalize-graph.ts transforms raw ASTs into the unified graph schema, ensuring consistent node types across languages.

  4. LLM Enrichment: llm-analyzer.ts formats graph nodes into prompts, calls the configured LLM (defaulting to GPT-4o-mini), and injects semantic metadata back into the graph.

  5. Persistence: The enriched graph is validated by schema.ts and persisted via persistence/index.ts, then consumed by packages/dashboard/src/store.ts for UI rendering.

The dashboard specifically imports only browser-safe sub-path exports (./search, ./types, ./schema) to prevent the heavyweight Tree-sitter WASM bundle from executing in the browser, keeping the UI lightweight and responsive.

Programmatic Usage Example

You can trigger the full pipeline programmatically using the analyzeProject function from the core package. This example requires Node.js >= 22 and assumes a PNPM workspace setup:

import { analyzeProject } from '@understand-anything/core';
import { writeFile } from 'fs/promises';

(async () => {
  // Point to the repository root (absolute path)
  const repoRoot = '/path/to/your/project';

  // Run the complete analysis pipeline:
  // discovery → tree-sitter parsing → graph building → LLM enrichment
  const graph = await analyzeProject(repoRoot, {
    llm: { model: 'gpt-4o-mini', temperature: 0.0 },
    languages: ['javascript', 'typescript', 'java'],
  });

  // Access enriched nodes with semantic metadata
  console.log('First node summary:', graph.nodes[0]?.summary);

  // Persist for dashboard consumption
  await writeFile(
    `${repoRoot}/.understand-anything/knowledge-graph.json`,
    JSON.stringify(graph, null, 2)
  );
})();

After building the core package and running the dashboard:

pnpm --filter @understand-anything/core build
pnpm --filter @understand-anything/dashboard dev

The dashboard will read the generated knowledge-graph.json, validate it against schema.ts, and render an interactive graph where each node displays the LLM-generated summary and semantic tags.

Summary

  • Tree-sitter integration in tree-sitter-plugin.ts provides fast, incremental parsing via WebAssembly, generating precise CSTs for multiple languages.
  • Graph construction in graph-builder.ts normalizes these trees into a unified, language-agnostic knowledge graph representing the full codebase.
  • LLM enrichment in llm-analyzer.ts adds semantic understanding through few-shot prompting, generating tags, summaries, and security classifications.
  • Incremental updates ensure only changed fragments are re-parsed and re-analyzed, optimizing both compute and token costs.
  • Browser-safe architecture keeps the heavy Tree-sitter WASM in the Node.js backend while serving lightweight data to the React dashboard.

Frequently Asked Questions

What role does Tree-sitter play in the Egonex AI pipeline?

Tree-sitter serves as the deterministic syntactic foundation. Implemented in packages/core/src/plugins/tree-sitter-plugin.ts using the Web-tree-sitter library, it parses source files into Concrete Syntax Trees that capture exact grammatical relationships, scopes, and imports. This guarantees syntactic fidelity that LLMs alone cannot provide, while its incremental parsing capability enables rapid updates when code changes.

How does the LLM analyzer enrich the knowledge graph?

The LLM analyzer in packages/core/src/analyzer/llm-analyzer.ts receives serialized graph fragments and uses few-shot prompts to infer high-level semantic concepts. It identifies design patterns, security-critical entry points, and component types, then augments graph nodes with tags, human-readable summaries, and confidence scores. This semantic layer transforms raw syntax into actionable insights like automated code tours and vulnerability detection.

Can the system handle incremental code updates efficiently?

Yes. Tree-sitter's incremental parser re-parses only changed file fragments rather than entire files, while the LLM analyzer receives batched prompts containing only modified sub-trees. This architecture minimizes both computational overhead and LLM token usage, enabling real-time updates to the knowledge graph as developers write code.

Is the dashboard browser-compatible with Tree-sitter?

The dashboard intentionally avoids running Tree-sitter in the browser. Instead, it imports only browser-safe sub-paths (./search, ./types, ./schema) and consumes pre-computed knowledge graphs persisted by the backend. The heavyweight Tree-sitter WASM bundle executes exclusively in the Node.js backend during the analysis phase, ensuring the React UI remains lightweight and responsive.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →