How Tree-Sitter-Based Extractors Work in Understand-Anything: Architecture and Registration

Tree-sitter-based extractors in Understand-Anything are registered via the TreeSitterPlugin class, which maps language IDs to extractor instances and transforms raw syntax tree nodes into normalized structural data.

The Egonex-AI/Understand-Anything repository leverages tree-sitter as its core parsing engine to perform structural analysis across multiple programming languages. Tree-sitter-based extractors implement a standardized LanguageExtractor interface to walk parse trees and extract functions, classes, imports, exports, and call-graph edges. This article explains the registration mechanism and the end-to-end flow from file path to structural analysis.

The TreeSitterPlugin Orchestrator

The TreeSitterPlugin class located in packages/core/src/plugins/tree-sitter-plugin.ts serves as the central orchestrator for all tree-sitter operations. It manages grammar loading, parser instantiation, and extractor dispatch.

Construction and Language Configuration

When instantiating the plugin, you can supply an optional array of LanguageConfig objects and an optional array of custom LanguageExtractor instances. If no configurations are provided, the constructor falls back to TypeScript and JavaScript grammars for backward compatibility.

constructor(configs?: LanguageConfig[], extractors?: LanguageExtractor[]) {
  // Builds _extensionToLang map from configs
  // Registers extractors if provided, otherwise uses builtins
}

The plugin builds an internal _extensionToLang map that associates file extensions (.ts, .js, .java, .go) with tree-sitter language IDs. This mapping enables constant-time language lookup based on file paths.

Extractor Registration

The registerExtractor method is the core mechanism for wiring extractors into the system. It accepts a LanguageExtractor instance and stores it in a private Map keyed by every language ID the extractor declares support for.

registerExtractor(extractor: LanguageExtractor): void {
  for (const id of extractor.languageIds) {
    this.extractors.set(id, extractor);
  }
}

If the constructor receives no custom extractors, it automatically imports and registers all built-in extractors exported from extractors/index.ts. This ensures that Java, Go, C#, C++, and TypeScript/JavaScript extractors are available by default.

Grammar Loading and Initialization

The init() method must be called before parsing. It dynamically loads the WASM binaries for each configured language using LanguageCls.load(wasmPath). All grammars are loaded asynchronously before any synchronous parsing occurs, ensuring parsers are ready for immediate use.

Parsing and Analysis Workflow

The analyzeFile method coordinates the end-to-end extraction process:

  1. Language Detection: Determines the language ID from the file extension using _extensionToLang.
  2. Parser Retrieval: getParser(filePath) creates a tree-sitter parser and sets its language.
  3. Tree Parsing: Parses the source content into a syntax tree.
  4. Extractor Lookup: getExtractor(langKey) retrieves the registered extractor, normalizing "tsx" to "typescript" for lookup purposes.
  5. Structure Extraction: Calls extractor.extractStructure(tree.rootNode) to obtain normalized data.

If no extractor exists for a language, analyzeFile returns empty arrays for functions, classes, imports, and exports.

The LanguageExtractor Interface

All extractors implement the LanguageExtractor interface defined in packages/core/src/plugins/extractors/types.ts. This contract ensures consistent behavior across language implementations.

Interface Contract

A minimal extractor must provide three properties:

  • languageIds: string[]: An array of language identifiers (e.g., ["java"], ["cpp", "c"]) that the extractor handles.
  • extractStructure(rootNode): StructuralAnalysis: Walks the tree-sitter tree and returns an object containing functions, classes, imports, and exports arrays.
  • extractCallGraph(rootNode): CallGraphEntry[]: Optionally extracts call-graph edges by identifying method invocations and mapping them to entries.

Language-Specific Implementations

The repository contains concrete implementations for multiple languages:

Each extractor walks the specific node types defined by its tree-sitter grammar and returns normalized data structures independent of the source language.

Built-in Extractor Registration

The Extractor Index

The extractors/index.ts file serves as the aggregation point for all built-in extractors. It imports individual extractor modules and exports them as a single array:

import { javaExtractor } from "./java-extractor.js";
import { goExtractor } from "./go-extractor.js";
import { csharpExtractor } from "./csharp-extractor.js";
import { cppExtractor } from "./cpp-extractor.js";
import { baseExtractor } from "./base-extractor.js";

export const builtinExtractors = [
  baseExtractor,
  javaExtractor,
  goExtractor,
  csharpExtractor,
  cppExtractor,
];

Automatic Registration Flow

When TreeSitterPlugin is constructed without custom extractors, it imports builtinExtractors and registers each one automatically:

for (const extractor of builtinExtractors) {
  this.registerExtractor(extractor);
}

Because registerExtractor maps every language ID declared by an extractor, the plugin can resolve the correct extractor in constant time via this.extractors.get(langKey).

End-to-End Usage Examples

Analyzing a Single File

You can use the plugin directly to analyze individual files:

import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";
import { readFile } from "node:fs/promises";

const plugin = new TreeSitterPlugin(); // Defaults to TS/JS
await plugin.init();

const source = await readFile("src/example.ts", "utf8");
const analysis = plugin.analyzeFile("src/example.ts", source);

console.log("Functions:", analysis.functions.map(f => f.name));
console.log("Imports:", analysis.imports.map(i => i.source));

Registering Custom Extractors at Runtime

To add support for a custom language or domain-specific language, pass an array of extractors to the constructor:

import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";

const myExtractor = {
  languageIds: ["mydsl"],
  extractStructure(root) {
    // Walk nodes and return normalized structure
    return { functions: [], classes: [], imports: [], exports: [] };
  },
  extractCallGraph(root) {
    return [];
  },
};

const plugin = new TreeSitterPlugin(undefined, [myExtractor]);
await plugin.init(); // Automatically registers myExtractor

Integration with the Graph Builder

In production usage, the plugin feeds data into higher-level analysis pipelines:

import { TreeSitterPlugin } from "@understand-anything/core/plugins/tree-sitter-plugin";

const plugin = new TreeSitterPlugin(languageConfigs);
await plugin.init();

const files = await collectProjectFiles();
for (const { path, content } of files) {
  const analysis = plugin.analyzeFile(path, content);
  // Feed analysis into graph builder
}

Summary

  • Tree-sitter-based extractors transform raw syntax trees into normalized structural data via the LanguageExtractor interface.
  • TreeSitterPlugin orchestrates grammar loading, parser management, and extractor dispatch from tree-sitter-plugin.ts.
  • Registration occurs through registerExtractor(), which maps each extractor to its declared languageIds in a private Map.
  • Built-in extractors are aggregated in extractors/index.ts and registered automatically when no custom extractors are provided.
  • Analysis flow moves from file path → language ID → parser → extractor → structured output, with fallback to empty results for unsupported languages.

Frequently Asked Questions

How does TreeSitterPlugin handle file extensions it does not recognize?

When analyzeFile is called on a file with an unrecognized extension, the plugin cannot determine a language ID from its internal _extensionToLang map. Consequently, getExtractor returns null, and analyzeFile returns empty arrays for all structural elements. You can avoid this by ensuring all target file extensions are mapped in the LanguageConfig objects passed to the constructor.

Can multiple extractors be registered for the same language ID?

No, the registration mechanism uses a Map where language IDs are keys. If you register a second extractor for a language ID already in the map, it will overwrite the previous entry. The registerExtractor method iterates over extractor.languageIds and calls this.extractors.set(id, extractor) for each, so the last extractor registered for a given ID wins.

What is the difference between LanguageConfig and LanguageExtractor?

LanguageConfig objects define metadata for tree-sitter grammars, including the file extension mappings and WASM binary paths needed to initialize parsers. LanguageExtractor instances contain the logic to walk the parsed syntax tree and extract meaningful structures (functions, classes, imports). The plugin uses configs to know how to parse a file, and extractors to know what to extract from the resulting tree.

How does the plugin normalize TypeScript and TSX handling?

The getExtractor method contains specific logic to normalize language keys: if the detected language is "tsx", it looks up the extractor using "typescript" instead. This allows the base TypeScript extractor to handle both standard TypeScript and TSX files without requiring separate extractor instances, as shown in the source: const key = langKey === "tsx" ? "typescript" : langKey;.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →