How Tree-Sitter and LLM Hybrid Analysis Works in Understand Anything
Understand Anything combines deterministic Tree-sitter AST parsing with probabilistic LLM reasoning to generate a hybrid knowledge graph that captures both exact code structure and semantic meaning.
The open-source Understand Anything project (Egonex-AI/Understand-Anything) employs a multi-phase pipeline that leverages the strengths of static analysis and artificial intelligence. By using Tree-sitter for precise, language-agnostic parsing and large language models for contextual interpretation, the system creates a rich representation of codebases that powers interactive exploration and automated analysis.
Architecture Overview
The hybrid analysis workflow operates in distinct phases, moving from raw source files to a unified knowledge graph. Each phase handles specific responsibilities while maintaining clean separation between structural and semantic concerns.
Phase 1: Language Configuration and Grammar Loading
The process begins in packages/core/src/plugins/tree-sitter-plugin.ts, where the TreeSitterPlugin class initializes the parsing infrastructure. During the init() method, the plugin loads WebAssembly grammar packages for each supported language defined in the language configuration files.
Each language entry specifies a treeSitter configuration pointing to WASM grammars (e.g., tree-sitter-typescript). The plugin constructs an extension-to-language mapping that enables rapid file type detection. This pre-loading step ensures grammars are ready before any file analysis begins.
Phase 2: File Parsing and Language Detection
When analyzing a specific file, the languageKeyFromPath() function resolves file extensions to language keys (e.g., ".tsx" → tsx). The plugin then instantiates a Tree-sitter Parser using the corresponding TreeSitterLanguage object.
The parser generates a concrete syntax tree (CST) that represents the exact grammatical structure of the source code. This deterministic parsing guarantees accurate extraction of functions, classes, imports, and exports regardless of code complexity or formatting.
Phase 3: Structural Extraction via Language Extractors
Once the AST is available, the plugin delegates to language-specific extractors located in packages/core/src/plugins/extractors/*.ts. Extractors implement the LanguageExtractor interface and translate Tree-sitter nodes into the unified StructuralAnalysis shape.
For example, typescript-extractor.ts traverses the TypeScript AST to identify callable entities, class hierarchies, and call-graph edges. The resulting structural data includes precise locations, signatures, and relationships derived directly from the grammar.
Phase 4: Fingerprinting and Deduplication
The fingerprint.ts module processes the StructuralAnalysis output to create a structural fingerprint. For supported languages, this fingerprint captures the AST-derived structure; for unsupported languages, it falls back to content hashing. This mechanism enables efficient change detection and incremental analysis across large repositories.
LLM Augmentation and Semantic Analysis
After structural extraction, the system enhances the deterministic data with LLM-powered semantic analysis. The llm-analyzer.ts module coordinates this phase by building structured prompts that combine raw source code with the structural context (function lists, class definitions, and call-graph entries).
The LLM receives this composite context and generates higher-level insights including file summaries, semantic tags, complexity ratings, and language-specific notes. The buildFileAnalysisPrompt() function constructs these prompts, while parseFileAnalysisResponse() handles the JSON response parsing and validation.
This augmentation adds probabilistic understanding that captures intent, architectural patterns, and cross-file relationships that static analysis alone cannot determine.
Knowledge Graph Assembly
The final phase occurs in packages/core/src/analyzer/graph-builder.ts, where structural and LLM outputs merge into cohesive knowledge graph nodes. Each node contains:
- Deterministic fields: Exact AST-derived entities (functions, classes, imports) from the Tree-sitter extraction
- Probabilistic fields: Semantic summaries, inferred relationships, and complexity metrics from the LLM
The graph-builder.ts module ensures these hybrid representations maintain referential integrity while supporting the dashboard's interactive navigation capabilities. Users can drill down using precise AST coordinates while benefiting from natural-language descriptions of code purpose.
Implementation Walkthrough
The following TypeScript example demonstrates the complete hybrid analysis pipeline:
// 1️⃣ Initialise the Tree-sitter plugin
import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";
import { defaultLanguageConfigs } from "./languages/configs/index.js";
const treeSitter = new TreeSitterPlugin(defaultLanguageConfigs);
await treeSitter.init(); // loads all WASM grammars into memory
// 2️⃣ Analyse a file structurally
const filePath = "src/app.ts";
const content = await readFile(filePath, "utf8");
const structural = await treeSitter.analyzeFile(filePath, content);
// 3️⃣ Pass structural data to the LLM analyser
import {
buildFileAnalysisPrompt,
parseFileAnalysisResponse
} from "./analyzer/llm-analyzer.js";
const projectCtx = "A TypeScript web app built with React and Express.";
const prompt = buildFileAnalysisPrompt(filePath, content, structural, projectCtx);
const llmResponse = await llmClient.generate(prompt);
const llmResult = parseFileAnalysisResponse(llmResponse);
// 4️⃣ Combine both results into a knowledge-graph node
const node = {
id: `file:${filePath}`,
type: "file",
path: filePath,
structural, // deterministic AST-derived data
...llmResult, // semantic summary, tags, complexity, etc.
};
graph.addNode(node);
Key Technical Components
| Component | Source File | Responsibility |
|---|---|---|
| Grammar Loading | packages/core/src/plugins/tree-sitter-plugin.ts |
Resolves WASM paths and pre-loads Tree-sitter grammars during init() |
| File Mapping | packages/core/src/plugins/tree-sitter-plugin.ts |
Maps file extensions to language keys via languageKeyFromPath() |
| AST Extraction | packages/core/src/plugins/extractors/*.ts |
Language-specific extractors implementing the LanguageExtractor interface |
| Fingerprinting | packages/core/src/fingerprint.ts |
Generates structural fingerprints or hash-only fallbacks |
| LLM Integration | packages/core/src/analyzer/llm-analyzer.ts |
Builds prompts and parses JSON responses for semantic analysis |
| Graph Assembly | packages/core/src/analyzer/graph-builder.ts |
Merges structural and LLM data into the final knowledge graph |
Summary
- Tree-sitter provides the foundation: Fast, deterministic parsing via WASM grammars ensures accurate extraction of code structure across multiple languages.
- Language extractors bridge the gap: The
LanguageExtractorinterface inpackages/core/src/plugins/extractors/translates AST nodes into a unifiedStructuralAnalysisformat. - LLMs add semantic depth: The
llm-analyzer.tsmodule enriches structural data with summaries, tags, and complexity ratings that require contextual understanding. - Hybrid nodes power exploration: The knowledge graph combines deterministic AST coordinates with probabilistic insights, enabling both precise navigation and natural-language querying.
Frequently Asked Questions
How does Understand Anything handle unsupported programming languages?
For unsupported languages, the system falls back to content hashing via fingerprint.ts instead of structural analysis. While the Tree-sitter plugin cannot extract AST data without a grammar, the LLM analyzer can still process the raw source code to generate basic summaries and tags, though with reduced precision compared to fully supported languages.
What is the performance impact of running both Tree-sitter and LLM analysis?
Tree-sitter parsing occurs locally and completes in milliseconds per file, making it suitable for real-time analysis. LLM augmentation introduces latency dependent on the provider's API response times, but the system mitigates this through structural fingerprinting that enables incremental updates—only changed files trigger new LLM requests.
How does the system ensure accuracy when merging LLM outputs with AST data?
The graph-builder.ts module maintains strict separation between deterministic structural fields (populated by Tree-sitter extractors) and probabilistic LLM fields. The parseFileAnalysisResponse() function validates JSON schemas and includes error handling for malformed LLM outputs, ensuring that structural integrity persists even when semantic analysis fails or hallucinates.
Can I extend the system to support a new programming language?
Yes, by implementing the LanguageExtractor interface in a new file within packages/core/src/plugins/extractors/ and adding the corresponding Tree-sitter grammar to the language configuration. The plugin architecture in tree-sitter-plugin.ts automatically discovers and loads new extractors during initialization, provided they follow the established pattern for AST traversal and StructuralAnalysis generation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →