How Tree-sitter and LLM Hybrid Static Analysis Works in Understand Anything

Understand Anything combines deterministic Tree-sitter parsing for precise AST extraction with LLM-driven semantic analysis to generate human-readable summaries and architectural insights.

The Egonex-AI/Understand-Anything repository implements a sophisticated hybrid analysis pipeline that merges deterministic code parsing with artificial intelligence interpretation. This Tree-sitter and LLM hybrid static analysis approach delivers accurate structural data alongside high-level semantic understanding of any codebase. By combining these complementary techniques, the tool creates a rich knowledge graph that powers interactive dashboards for code navigation.

The Hybrid Architecture: Precision Meets Interpretation

The analysis pipeline operates as a two-stage system where deterministic parsing feeds into interpretive reasoning.

Stage 1: Tree-sitter Structural Extraction

The first stage relies on Tree-sitter, a fast, grammar-driven parser that generates a precise abstract syntax tree (AST) for every supported language. According to the source code in packages/core/src/plugins/tree-sitter-plugin.ts, the plugin instantiates with a list of LanguageConfig objects or falls back to built-in TypeScript/JavaScript support.

On init(), the system loads WebAssembly grammars for each configured language, including optional TSX support. The analyzeFile() method creates a language-specific parser (determined from file extension mappings), parses the source, and delegates the AST root node to a language-specific extractor. This extractor returns a StructuralAnalysis object enumerating functions, classes, imports, exports, and optionally a call-graph.

Stage 2: LLM Semantic Enrichment

The second stage, implemented in packages/core/src/analyzer/llm-analyzer.ts, uses large language models to interpret the structural data. Helper functions buildFileAnalysisPrompt() and buildProjectSummaryPrompt() craft prompts that embed the raw file contents alongside project context, requesting strict JSON output. The analyser then uses parseFileAnalysisResponse() and parseProjectSummaryResponse() to safely extract and normalize the JSON into strongly-typed objects.

Deep Dive: Tree-sitter Plugin Implementation

The Tree-sitter plugin serves as the deterministic foundation of the pipeline. When processing a file, the system:

  1. Determines the language from the file extension using an internal mapping
  2. Loads the appropriate WebAssembly grammar (supporting languages like TypeScript, JavaScript, and TSX)
  3. Parses the source into an AST
  4. Extracts structural elements through language-specific extractors

The StructuralAnalysis object contains ground-truth syntactic facts including line numbers, identifier names, and module relationships. This data provides the reliable foundation upon which the LLM builds its interpretations.

LLM Analyzer: From Structure to Meaning

While Tree-sitter provides the "what," the LLM provides the "why" and "how." The buildFileAnalysisPrompt() function constructs prompts that include:

  • The raw source code content
  • The extracted structural data (function names, import sources, class hierarchies)
  • Project-wide context describing the codebase purpose

The LLM returns JSON containing file summaries, architectural tags, complexity ratings, and layer descriptions (e.g., "this file implements a REST API" or "UI → API → Data layers"). The parseFileAnalysisResponse() function handles validation and type conversion, ensuring the output integrates cleanly into the knowledge graph.

The Complete Hybrid Workflow

The full pipeline, orchestrated through components like packages/core/src/analyzer/graph-builder.ts, follows four distinct steps:

  1. Static extraction – For each source file, the Tree-sitter plugin yields a StructuralAnalysis containing precise AST data
  2. Prompt enrichment – The analyser builds LLM prompts combining the extracted structure with raw source code and project context
  3. LLM inference – The language model processes the enriched prompt and returns JSON summaries (file-level or project-level)
  4. Post-processing – Response parsers normalize the JSON into typed objects and merge them into the knowledge graph

If a language lacks a Tree-sitter grammar, the system gracefully falls back to pure code-only analysis, ensuring complete repository coverage regardless of language support.

Why This Hybrid Approach Works

This architecture delivers three key advantages:

  • Precision – Tree-sitter guarantees accurate syntactic data (line numbers, identifier names) without heuristics or guesswork
  • Expressiveness – The LLM reasons about intent, architectural patterns, and complexity that are not directly inferable from the AST alone
  • Scalability – Parsing operates in O(N) time relative to file size, while LLM calls are limited to representative samples, keeping costs modest

Summary

  • Tree-sitter and LLM hybrid static analysis combines deterministic AST parsing with AI-driven semantic interpretation
  • The TreeSitterPlugin class in packages/core/src/plugins/tree-sitter-plugin.ts handles language detection, WebAssembly grammar loading, and structural extraction
  • The LLM analyzer in packages/core/src/analyzer/llm-analyzer.ts provides prompt builders (buildFileAnalysisPrompt, buildProjectSummaryPrompt) and response parsers (parseFileAnalysisResponse, parseProjectSummaryResponse)
  • The system gracefully degrades to pure LLM analysis when Tree-sitter grammars are unavailable for specific languages
  • Extracted structural data includes functions, classes, imports/exports, and call-graphs, enabling rich code navigation

Frequently Asked Questions

What happens if Tree-sitter doesn't support a specific programming language?

If a language lacks a Tree-sitter grammar, the system falls back to pure LLM-driven analysis using only the raw source code. While this loses the precise structural guarantees of AST parsing, it ensures every file in the repository receives some level of analysis and inclusion in the knowledge graph.

How does the Tree-sitter plugin determine which grammar to use for a file?

The plugin uses an extension-to-language mapping configured during initialization. When analyzeFile() receives a file path, it extracts the extension and matches it against the configured LanguageConfig objects. If no match exists, it defaults to the built-in TypeScript/JavaScript grammars or returns an error for unsupported languages.

What specific structural data does the Tree-sitter extraction capture?

The extractor returns a StructuralAnalysis object containing function declarations, class definitions, import and export statements, and optionally a call-graph showing relationships between functions. This data includes precise line numbers and identifier names, providing ground-truth coordinates for the LLM's semantic analysis.

How does the system manage costs when analyzing large repositories?

The architecture limits LLM calls to representative file samples rather than processing every file through the language model. Tree-sitter parsing handles all files locally at O(N) speed, while the LLM processes only strategically selected files or aggregated project summaries. This hybrid approach keeps API costs modest while maintaining comprehensive coverage through deterministic parsing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →