# How the Tree-Sitter Plugin Extracts Structural Facts from Source Code in Understand-Anything

> Discover how the Tree-Sitter plugin transforms source code into a language-agnostic description of project structure in Understand-Anything by parsing ASTs and using language-specific extractors.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: internals
- Published: 2026-06-18

---

**The `TreeSitterPlugin` transforms raw source files into a rich, language-agnostic description of project structure by parsing ASTs with Web-Tree-Sitter and delegating to language-specific extractors.**

The **Egonex-AI/Understand-Anything** repository relies on a robust static analysis pipeline to build knowledge graphs from source code. At the heart of this pipeline lies the tree-sitter plugin, which bridges raw text files and structured data by leveraging WebAssembly grammars and language-specific extractors.

## Configuration-Driven Grammar Loading

The plugin initializes through a declarative configuration system. During construction, it receives a list of `LanguageConfig` objects, each specifying a language identifier, file extensions, and the location of a Web-Tree-Sitter WASM grammar defined by `treeSitter.wasmPackage` and `wasmFile` properties.

In [`packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/tree-sitter-plugin.ts) (lines 20-38), the `init()` method pre-loads every grammar into memory. This design ensures that subsequent parsing operations remain completely synchronous, eliminating async overhead during the critical analysis phase.

## Extension-to-Language Mapping

Before processing files, the constructor builds an internal map from file extensions (e.g., `.ts`, `.js`, `.tsx`) to their corresponding language IDs. This mapping, implemented in lines 64-83 of [`tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/tree-sitter-plugin.ts), enables the plugin to instantly select the correct grammar when encountering any source file.

## Synchronous Parser Creation

The `getParser(filePath)` method (lines 99-119) instantiates a new `web-tree-sitter` `Parser` object, configures it with the language matching the file extension, and returns it ready for use. If no grammar exists for the given extension, the method returns `null`, allowing the broader pipeline to execute graceful fallback logic rather than throwing exceptions.

## Structural Analysis and the Extractor Pattern

The `analyzeFile(filePath, content)` method (lines 121-150 in [`tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/tree-sitter-plugin.ts)) serves as the primary entry point for fact extraction. It parses the source into an AST, retrieves the appropriate language-specific extractor (such as `TypeScriptExtractor` or `PythonExtractor`), and invokes `extractStructure(rootNode)`.

Walking the Tree-Sitter node tree, the extractor produces a `StructuralAnalysis` object containing:

- **functions** – name, parameters, return type, and source line range
- **classes** – name, methods, properties, and line range
- **imports** – source string, imported identifiers, and line number
- **exports** – exported names including default and aliased exports

Language-specific extractors reside in [`packages/core/src/plugins/extractors/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/extractors/index.ts) (lines 27-40), which registers built-in extractors for TypeScript, Python, Go, and other supported languages.

## Practical Usage Example

```typescript
import { TreeSitterPlugin } from '@understand-anything/core';

// 1️⃣ Create the plugin (no explicit configs → defaults to TS/JS)
const plugin = new TreeSitterPlugin();
await plugin.init();               // load WASM grammars

// 2️⃣ Analyse a TypeScript file
const tsSource = `
  export class Service {
    private id: number;
    constructor(id: number) { this.id = id; }
    greet(name: string) { return \`Hi \${name}\`; }
  }
`;
const analysis = plugin.analyzeFile('service.ts', tsSource);
console.log(analysis.classes[0].methods); // → [ 'constructor', 'greet' ]

// 3️⃣ Resolve imports in a file
const jsSource = `
  import { foo } from './utils';
  import * as path from 'path';
`;
const resolved = plugin.resolveImports('/proj/src/app.js', jsSource);
console.log(resolved);
/* → [
     { source: './utils', resolvedPath: '/proj/src/utils', specifiers: ['foo'] },
     { source: 'path',    resolvedPath: 'path',          specifiers: ['* as path'] }
   ] */

```

## Import Resolution and Call Graph Extraction

Beyond static structure, the plugin enables cross-file analysis through two specialized methods.

**Resolving Relative Imports**

The `resolveImports()` method (lines 152-176) reuses the structural analysis to convert relative import strings like `./foo` into absolute file system paths, while preserving external package names (such as `path` or `react`) unchanged. This allows downstream graph builders to distinguish between internal project dependencies and external libraries.

**Building Call Graphs**

The `extractCallGraph()` method (lines 178-197) returns a list of `CallGraphEntry` objects representing `caller → callee` relationships. By parsing the file and delegating to the language-specific extractor, it identifies function invocations that the higher-level graph builder uses to connect functions across multiple files.

## Extending Support with Built-in Extractors

The plugin's extensibility stems from its extractor registry in [`packages/core/src/plugins/extractors/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/extractors/index.ts). Each extractor implements the logic for converting Tree-Sitter nodes into standardized structural facts. For example, [`packages/core/src/plugins/extractors/typescript-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/extractors/typescript-extractor.ts) demonstrates how to walk TypeScript-specific AST nodes to identify classes, methods, and import declarations.

The repository also includes compiled WASM grammars, such as `understand-anything-plugin/packages/tree-sitter-dart-wasm/tree-sitter-dart.wasm`, which can be activated by providing a corresponding `LanguageConfig` when instantiating the plugin.

## Summary

- The **TreeSitterPlugin** in [`packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/tree-sitter-plugin.ts) acts as the core static analysis engine for the Understand-Anything project.
- It pre-loads Web-Tree-Sitter WASM grammars during `init()` to ensure synchronous parsing performance.
- The `analyzeFile()` method delegates to language-specific extractors to produce structured data about functions, classes, imports, and exports.
- `resolveImports()` maps relative paths to absolute locations, while `extractCallGraph()` builds caller-to-callee relationships for dependency analysis.
- The architecture supports extension through the extractor registry in [`packages/core/src/plugins/extractors/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/extractors/index.ts).

## Frequently Asked Questions

### What file formats does the tree-sitter plugin support?

The plugin supports any language with a valid Web-Tree-Sitter WASM grammar and a registered extractor. By default, it includes built-in extractors for TypeScript, JavaScript, Python, and Go, with the ability to add custom languages by providing a `LanguageConfig` pointing to a WASM file (such as the Dart example bundled in the repository).

### How does the plugin handle unsupported file extensions?

When `getParser(filePath)` encounters an extension without a mapped grammar, it returns `null` rather than throwing an error. This allows the broader analysis pipeline to skip unsupported files or apply alternative parsing strategies without crashing the extraction process.

### What is the difference between structural analysis and call graph extraction?

Structural analysis via `analyzeFile()` captures static declarations—functions, classes, imports, and exports—returning a `StructuralAnalysis` object. Call graph extraction via `extractCallGraph()` specifically identifies dynamic relationships between these entities, returning `CallGraphEntry` objects that map which functions call other functions, enabling cross-file dependency tracing.

### Where are the tree-sitter grammars stored and loaded from?

Grammars are loaded from locations specified in `LanguageConfig` objects passed to the plugin constructor. Each config defines `treeSitter.wasmPackage` and `wasmFile` properties pointing to WebAssembly binaries (e.g., `tree-sitter-dart.wasm`). The plugin loads these into memory during the asynchronous `init()` phase, making them available for synchronous parsing throughout the analysis lifecycle.