How Tree-sitter and LLM Hybrid Static Analysis Works in Understand Anything
Understand Anything combines deterministic Tree-sitter parsing for precise AST extraction with LLM-driven semantic analysis to generate human-readable summaries and architectural insights.
The Egonex-AI/Understand-Anything repository implements a sophisticated hybrid analysis pipeline that merges deterministic code parsing with artificial intelligence interpretation. This Tree-sitter and LLM hybrid static analysis approach delivers accurate structural data alongside high-level semantic understanding of any codebase. By combining these complementary techniques, the tool creates a rich knowledge graph that powers interactive dashboards for code navigation.
The Hybrid Architecture: Precision Meets Interpretation
The analysis pipeline operates as a two-stage system where deterministic parsing feeds into interpretive reasoning.
Stage 1: Tree-sitter Structural Extraction
The first stage relies on Tree-sitter, a fast, grammar-driven parser that generates a precise abstract syntax tree (AST) for every supported language. According to the source code in packages/core/src/plugins/tree-sitter-plugin.ts, the plugin instantiates with a list of LanguageConfig objects or falls back to built-in TypeScript/JavaScript support.
On init(), the system loads WebAssembly grammars for each configured language, including optional TSX support. The analyzeFile() method creates a language-specific parser (determined from file extension mappings), parses the source, and delegates the AST root node to a language-specific extractor. This extractor returns a StructuralAnalysis object enumerating functions, classes, imports, exports, and optionally a call-graph.
Stage 2: LLM Semantic Enrichment
The second stage, implemented in packages/core/src/analyzer/llm-analyzer.ts, uses large language models to interpret the structural data. Helper functions buildFileAnalysisPrompt() and buildProjectSummaryPrompt() craft prompts that embed the raw file contents alongside project context, requesting strict JSON output. The analyser then uses parseFileAnalysisResponse() and parseProjectSummaryResponse() to safely extract and normalize the JSON into strongly-typed objects.
Deep Dive: Tree-sitter Plugin Implementation
The Tree-sitter plugin serves as the deterministic foundation of the pipeline. When processing a file, the system:
- Determines the language from the file extension using an internal mapping
- Loads the appropriate WebAssembly grammar (supporting languages like TypeScript, JavaScript, and TSX)
- Parses the source into an AST
- Extracts structural elements through language-specific extractors
The StructuralAnalysis object contains ground-truth syntactic facts including line numbers, identifier names, and module relationships. This data provides the reliable foundation upon which the LLM builds its interpretations.
LLM Analyzer: From Structure to Meaning
While Tree-sitter provides the "what," the LLM provides the "why" and "how." The buildFileAnalysisPrompt() function constructs prompts that include:
- The raw source code content
- The extracted structural data (function names, import sources, class hierarchies)
- Project-wide context describing the codebase purpose
The LLM returns JSON containing file summaries, architectural tags, complexity ratings, and layer descriptions (e.g., "this file implements a REST API" or "UI → API → Data layers"). The parseFileAnalysisResponse() function handles validation and type conversion, ensuring the output integrates cleanly into the knowledge graph.
The Complete Hybrid Workflow
The full pipeline, orchestrated through components like packages/core/src/analyzer/graph-builder.ts, follows four distinct steps:
- Static extraction – For each source file, the Tree-sitter plugin yields a
StructuralAnalysiscontaining precise AST data - Prompt enrichment – The analyser builds LLM prompts combining the extracted structure with raw source code and project context
- LLM inference – The language model processes the enriched prompt and returns JSON summaries (file-level or project-level)
- Post-processing – Response parsers normalize the JSON into typed objects and merge them into the knowledge graph
If a language lacks a Tree-sitter grammar, the system gracefully falls back to pure code-only analysis, ensuring complete repository coverage regardless of language support.
Why This Hybrid Approach Works
This architecture delivers three key advantages:
- Precision – Tree-sitter guarantees accurate syntactic data (line numbers, identifier names) without heuristics or guesswork
- Expressiveness – The LLM reasons about intent, architectural patterns, and complexity that are not directly inferable from the AST alone
- Scalability – Parsing operates in O(N) time relative to file size, while LLM calls are limited to representative samples, keeping costs modest
Summary
- Tree-sitter and LLM hybrid static analysis combines deterministic AST parsing with AI-driven semantic interpretation
- The
TreeSitterPluginclass inpackages/core/src/plugins/tree-sitter-plugin.tshandles language detection, WebAssembly grammar loading, and structural extraction - The LLM analyzer in
packages/core/src/analyzer/llm-analyzer.tsprovides prompt builders (buildFileAnalysisPrompt,buildProjectSummaryPrompt) and response parsers (parseFileAnalysisResponse,parseProjectSummaryResponse) - The system gracefully degrades to pure LLM analysis when Tree-sitter grammars are unavailable for specific languages
- Extracted structural data includes functions, classes, imports/exports, and call-graphs, enabling rich code navigation
Frequently Asked Questions
What happens if Tree-sitter doesn't support a specific programming language?
If a language lacks a Tree-sitter grammar, the system falls back to pure LLM-driven analysis using only the raw source code. While this loses the precise structural guarantees of AST parsing, it ensures every file in the repository receives some level of analysis and inclusion in the knowledge graph.
How does the Tree-sitter plugin determine which grammar to use for a file?
The plugin uses an extension-to-language mapping configured during initialization. When analyzeFile() receives a file path, it extracts the extension and matches it against the configured LanguageConfig objects. If no match exists, it defaults to the built-in TypeScript/JavaScript grammars or returns an error for unsupported languages.
What specific structural data does the Tree-sitter extraction capture?
The extractor returns a StructuralAnalysis object containing function declarations, class definitions, import and export statements, and optionally a call-graph showing relationships between functions. This data includes precise line numbers and identifier names, providing ground-truth coordinates for the LLM's semantic analysis.
How does the system manage costs when analyzing large repositories?
The architecture limits LLM calls to representative file samples rather than processing every file through the language model. Tree-sitter parsing handles all files locally at O(N) speed, while the LLM processes only strategically selected files or aggregated project summaries. This hybrid approach keeps API costs modest while maintaining comprehensive coverage through deterministic parsing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →