# How Tree-Sitter-Based Extractors Work in Understand-Anything: Architecture and Registration

> Discover how tree-sitter-based extractors work in Understand-Anything. Learn about Egonex-AI's registration process and data transformation for structural analysis.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: architecture
- Published: 2026-06-14

---

**Tree-sitter-based extractors in Understand-Anything are registered via the `TreeSitterPlugin` class, which maps language IDs to extractor instances and transforms raw syntax tree nodes into normalized structural data.**

The Egonex-AI/Understand-Anything repository leverages tree-sitter as its core parsing engine to perform structural analysis across multiple programming languages. Tree-sitter-based extractors implement a standardized `LanguageExtractor` interface to walk parse trees and extract functions, classes, imports, exports, and call-graph edges. This article explains the registration mechanism and the end-to-end flow from file path to structural analysis.

## The TreeSitterPlugin Orchestrator

The `TreeSitterPlugin` class located in [`packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/tree-sitter-plugin.ts) serves as the central orchestrator for all tree-sitter operations. It manages grammar loading, parser instantiation, and extractor dispatch.

### Construction and Language Configuration

When instantiating the plugin, you can supply an optional array of `LanguageConfig` objects and an optional array of custom `LanguageExtractor` instances. If no configurations are provided, the constructor falls back to TypeScript and JavaScript grammars for backward compatibility.

```typescript
constructor(configs?: LanguageConfig[], extractors?: LanguageExtractor[]) {
  // Builds _extensionToLang map from configs
  // Registers extractors if provided, otherwise uses builtins
}

```

The plugin builds an internal `_extensionToLang` map that associates file extensions (`.ts`, `.js`, `.java`, `.go`) with tree-sitter language IDs. This mapping enables constant-time language lookup based on file paths.

### Extractor Registration

The `registerExtractor` method is the core mechanism for wiring extractors into the system. It accepts a `LanguageExtractor` instance and stores it in a private `Map` keyed by every language ID the extractor declares support for.

```typescript
registerExtractor(extractor: LanguageExtractor): void {
  for (const id of extractor.languageIds) {
    this.extractors.set(id, extractor);
  }
}

```

If the constructor receives no custom extractors, it automatically imports and registers all built-in extractors exported from [`extractors/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extractors/index.ts). This ensures that Java, Go, C#, C++, and TypeScript/JavaScript extractors are available by default.

### Grammar Loading and Initialization

The `init()` method must be called before parsing. It dynamically loads the WASM binaries for each configured language using `LanguageCls.load(wasmPath)`. All grammars are loaded asynchronously before any synchronous parsing occurs, ensuring parsers are ready for immediate use.

### Parsing and Analysis Workflow

The `analyzeFile` method coordinates the end-to-end extraction process:

1. **Language Detection**: Determines the language ID from the file extension using `_extensionToLang`.
2. **Parser Retrieval**: `getParser(filePath)` creates a tree-sitter parser and sets its language.
3. **Tree Parsing**: Parses the source content into a syntax tree.
4. **Extractor Lookup**: `getExtractor(langKey)` retrieves the registered extractor, normalizing `"tsx"` to `"typescript"` for lookup purposes.
5. **Structure Extraction**: Calls `extractor.extractStructure(tree.rootNode)` to obtain normalized data.

If no extractor exists for a language, `analyzeFile` returns empty arrays for functions, classes, imports, and exports.

## The LanguageExtractor Interface

All extractors implement the `LanguageExtractor` interface defined in [`packages/core/src/plugins/extractors/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/plugins/extractors/types.ts). This contract ensures consistent behavior across language implementations.

### Interface Contract

A minimal extractor must provide three properties:

- **`languageIds: string[]`**: An array of language identifiers (e.g., `["java"]`, `["cpp", "c"]`) that the extractor handles.
- **`extractStructure(rootNode): StructuralAnalysis`**: Walks the tree-sitter tree and returns an object containing `functions`, `classes`, `imports`, and `exports` arrays.
- **`extractCallGraph(rootNode): CallGraphEntry[]`**: Optionally extracts call-graph edges by identifying method invocations and mapping them to entries.

### Language-Specific Implementations

The repository contains concrete implementations for multiple languages:

- **Java**: [`java-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/java-extractor.ts) handles `class_declaration`, `method_declaration`, and `import_declaration` nodes.
- **Go**: [`go-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/go-extractor.ts) processes Go-specific syntax for packages and functions.
- **C#**: [`csharp-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/csharp-extractor.ts) analyzes C# compilation units.

- **C/C++**: [`cpp-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/cpp-extractor.ts) supports both `"cpp"` and `"c"` language IDs.
- **TypeScript/JavaScript**: [`base-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/base-extractor.ts) provides shared logic for `"typescript"`, `"tsx"`, and `"javascript"`.

Each extractor walks the specific node types defined by its tree-sitter grammar and returns normalized data structures independent of the source language.

## Built-in Extractor Registration

### The Extractor Index

The [`extractors/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extractors/index.ts) file serves as the aggregation point for all built-in extractors. It imports individual extractor modules and exports them as a single array:

```typescript
import { javaExtractor } from "./java-extractor.js";
import { goExtractor } from "./go-extractor.js";
import { csharpExtractor } from "./csharp-extractor.js";
import { cppExtractor } from "./cpp-extractor.js";
import { baseExtractor } from "./base-extractor.js";

export const builtinExtractors = [
  baseExtractor,
  javaExtractor,
  goExtractor,
  csharpExtractor,
  cppExtractor,
];

```

### Automatic Registration Flow

When `TreeSitterPlugin` is constructed without custom extractors, it imports `builtinExtractors` and registers each one automatically:

```typescript
for (const extractor of builtinExtractors) {
  this.registerExtractor(extractor);
}

```

Because `registerExtractor` maps **every** language ID declared by an extractor, the plugin can resolve the correct extractor in constant time via `this.extractors.get(langKey)`.

## End-to-End Usage Examples

### Analyzing a Single File

You can use the plugin directly to analyze individual files:

```typescript
import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";
import { readFile } from "node:fs/promises";

const plugin = new TreeSitterPlugin(); // Defaults to TS/JS
await plugin.init();

const source = await readFile("src/example.ts", "utf8");
const analysis = plugin.analyzeFile("src/example.ts", source);

console.log("Functions:", analysis.functions.map(f => f.name));
console.log("Imports:", analysis.imports.map(i => i.source));

```

### Registering Custom Extractors at Runtime

To add support for a custom language or domain-specific language, pass an array of extractors to the constructor:

```typescript
import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";

const myExtractor = {
  languageIds: ["mydsl"],
  extractStructure(root) {
    // Walk nodes and return normalized structure
    return { functions: [], classes: [], imports: [], exports: [] };
  },
  extractCallGraph(root) {
    return [];
  },
};

const plugin = new TreeSitterPlugin(undefined, [myExtractor]);
await plugin.init(); // Automatically registers myExtractor

```

### Integration with the Graph Builder

In production usage, the plugin feeds data into higher-level analysis pipelines:

```typescript
import { TreeSitterPlugin } from "@understand-anything/core/plugins/tree-sitter-plugin";

const plugin = new TreeSitterPlugin(languageConfigs);
await plugin.init();

const files = await collectProjectFiles();
for (const { path, content } of files) {
  const analysis = plugin.analyzeFile(path, content);
  // Feed analysis into graph builder
}

```

## Summary

- **Tree-sitter-based extractors** transform raw syntax trees into normalized structural data via the `LanguageExtractor` interface.
- **`TreeSitterPlugin`** orchestrates grammar loading, parser management, and extractor dispatch from [`tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/tree-sitter-plugin.ts).
- **Registration** occurs through `registerExtractor()`, which maps each extractor to its declared `languageIds` in a private `Map`.
- **Built-in extractors** are aggregated in [`extractors/index.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/extractors/index.ts) and registered automatically when no custom extractors are provided.
- **Analysis flow** moves from file path → language ID → parser → extractor → structured output, with fallback to empty results for unsupported languages.

## Frequently Asked Questions

### How does TreeSitterPlugin handle file extensions it does not recognize?

When `analyzeFile` is called on a file with an unrecognized extension, the plugin cannot determine a language ID from its internal `_extensionToLang` map. Consequently, `getExtractor` returns `null`, and `analyzeFile` returns empty arrays for all structural elements. You can avoid this by ensuring all target file extensions are mapped in the `LanguageConfig` objects passed to the constructor.

### Can multiple extractors be registered for the same language ID?

No, the registration mechanism uses a `Map` where language IDs are keys. If you register a second extractor for a language ID already in the map, it will overwrite the previous entry. The `registerExtractor` method iterates over `extractor.languageIds` and calls `this.extractors.set(id, extractor)` for each, so the last extractor registered for a given ID wins.

### What is the difference between `LanguageConfig` and `LanguageExtractor`?

`LanguageConfig` objects define metadata for tree-sitter grammars, including the file extension mappings and WASM binary paths needed to initialize parsers. `LanguageExtractor` instances contain the logic to walk the parsed syntax tree and extract meaningful structures (functions, classes, imports). The plugin uses configs to know *how* to parse a file, and extractors to know *what* to extract from the resulting tree.

### How does the plugin normalize TypeScript and TSX handling?

The `getExtractor` method contains specific logic to normalize language keys: if the detected language is `"tsx"`, it looks up the extractor using `"typescript"` instead. This allows the base TypeScript extractor to handle both standard TypeScript and TSX files without requiring separate extractor instances, as shown in the source: `const key = langKey === "tsx" ? "typescript" : langKey;`.