How to Extend Understand-Anything to Support a New Programming Language Using base-extractor
To extend Understand-Anything with a new language, create a class implementing the LanguageExtractor interface under packages/core/src/plugins/extractors/ that imports traversal utilities from base-extractor.ts, export it from the extractors index, register a language configuration in packages/core/src/languages/configs/ declaring the Tree-Sitter grammar and file extensions, and run the core test suite to verify integration.
Understand-Anything is an extensible static analysis framework that leverages Tree-Sitter grammars to parse multiple programming languages. The core architecture relies on the shared utility module base-extractor.ts to provide common AST traversal operations, allowing you to add support for additional languages without duplicating logic. By implementing the LanguageExtractor interface and registering your implementation with the plugin system, you enable structural analysis and call-graph extraction for any Tree-Sitter-supported language.
Overview of the Extension Pipeline
Extending the framework follows a four-phase pipeline that isolates language-specific logic while maximizing code reuse:
- Create a language-specific extractor implementing the
LanguageExtractorinterface inpackages/core/src/plugins/extractors/, importing helpers frombase-extractor.ts. - Register the extractor by exporting it from
packages/core/src/plugins/extractors/index.tsso theTreeSitterPlugincan discover it. - Add a language configuration in
packages/core/src/languages/configs/that defines the language ID, file extensions, and Tree-Sitter grammar dependency. - Wire everything into the plugin by adding the config to
builtinLanguageConfigsand running the test suite to verify integration.
The TreeSitterPlugin automatically coordinates these components, routing parsed ASTs to the correct extractor based on the language configuration.
Step 1: Create the Language-Specific Extractor
Create a new file under packages/core/src/plugins/extractors/ (for example, my-lang-extractor.ts) that implements the LanguageExtractor interface. This class must define languageIds and implement two methods: extractStructure and extractCallGraph.
Import the shared traversal utilities from base-extractor.ts to handle common AST operations:
// 📄 packages/core/src/plugins/extractors/my-lang-extractor.ts
import type { StructuralAnalysis, CallGraphEntry } from "../../types.js";
import type { LanguageExtractor, TreeSitterNode } from "./types.js";
import {
traverse,
getStringValue,
findChild,
findChildren,
hasChildOfType,
} from "./base-extractor.js";
export class MyLangExtractor implements LanguageExtractor {
/** Language IDs must match the LanguageConfig.id */
readonly languageIds = ["my-lang"];
/** Structural analysis – functions, classes, imports, exports */
extractStructure(root: TreeSitterNode): StructuralAnalysis {
const functions: StructuralAnalysis["functions"] = [];
const classes: StructuralAnalysis["classes"] = [];
const imports: StructuralAnalysis["imports"] = [];
const exports: StructuralAnalysis["exports"] = [];
traverse(root, (node) => {
switch (node.type) {
case "function_declaration":
functions.push({
name: findChild(node, "identifier")!.text,
lineRange: [
node.startPosition.row + 1,
node.endPosition.row + 1,
],
params: [],
returnType: undefined,
});
break;
case "class_declaration":
classes.push({
name: findChild(node, "type_identifier")!.text,
lineRange: [
node.startPosition.row + 1,
node.endPosition.row + 1,
],
methods: [],
properties: [],
});
break;
case "import_statement":
imports.push({
source: getStringValue(findChild(node, "string")!),
specifiers: [],
lineNumber: node.startPosition.row + 1,
});
break;
case "export_statement":
exports.push({
name: findChild(node, "identifier")!.text,
lineNumber: node.startPosition.row + 1,
isDefault: hasChildOfType(node, "default"),
});
break;
}
});
return { functions, classes, imports, exports };
}
/** Call-graph extraction – simple caller → callee mapping */
extractCallGraph(root: TreeSitterNode): CallGraphEntry[] {
const entries: CallGraphEntry[] = [];
const stack: string[] = [];
const walk = (node: TreeSitterNode) => {
if (node.type === "function_declaration") {
const name = findChild(node, "identifier")?.text;
if (name) stack.push(name);
}
if (node.type === "call_expression") {
const callee = findChild(node, "identifier")?.text;
if (callee && stack.length) {
entries.push({
caller: stack[stack.length - 1],
callee,
lineNumber: node.startPosition.row + 1,
});
}
}
for (let i = 0; i < node.childCount; i++) {
const child = node.child(i);
if (child) walk(child);
}
if (node.type === "function_declaration") stack.pop();
};
walk(root);
return entries;
}
}
Key implementation details:
- Import helpers from
base-extractor.ts: The functionstraverse,getStringValue,findChild,findChildren, andhasChildOfTypeprovide standard AST navigation without manual recursion boilerplate. - Match
languageIdsto config: ThelanguageIdsarray must contain the same identifier defined in the language configuration file (e.g.,"my-lang"). - Use grammar-specific node types: Strings like
"function_declaration"and"class_declaration"correspond to node types emitted by the Tree-Sitter grammar for your target language.
Step 2: Register the Extractor
Export the new extractor class from the central index so the TreeSitterPlugin can locate it dynamically:
// 📄 packages/core/src/plugins/extractors/index.ts
export {
TypeScriptExtractor,
// …other existing extractors
MyLangExtractor,
} from "./my-lang-extractor.js";
The plugin system discovers extractors through this export list and instantiates the appropriate class based on the file's detected language ID.
Step 3: Add a Language Configuration
Create a configuration file in packages/core/src/languages/configs/ that declares the language metadata and Tree-Sitter grammar:
// 📄 packages/core/src/languages/configs/my-lang.ts
/**
* Language configuration for MyLang
* Add the Tree-Sitter grammar (e.g., tree-sitter-my-lang) to package.json dependencies.
*/
export const myLangConfig = {
id: "my-lang",
name: "MyLang",
extensions: [".my"],
treeSitter: {
language: "my-lang",
grammar: "tree-sitter-my-lang", // npm package name
version: "^0.0.1",
},
};
Then register this configuration in the built-in languages index:
// 📄 packages/core/src/languages/configs/index.ts
import { myLangConfig } from "./my-lang.js";
export const builtinLanguageConfigs = [
// ...existing configs
myLangConfig,
];
Important: Add the Tree-Sitter grammar package (e.g., tree-sitter-my-lang) to packages/core/package.json under dependencies and run pnpm install to make the parser available at runtime.
Step 4: Test the New Extractor
Verify your implementation by creating a dedicated test file following the existing patterns:
// 📄 packages/core/src/plugins/extractors/__tests__/my-lang-extractor.test.ts
import { MyLangExtractor } from "../my-lang-extractor.js";
test("MyLangExtractor reports correct languageIds", () => {
const extractor = new MyLangExtractor();
expect(extractor.languageIds).toEqual(["my-lang"]);
});
test("MyLangExtractor extracts function declarations", () => {
// Add test cases using a parsed AST fixture
});
Run the core test suite to confirm integration:
pnpm --filter @understand-anything/core test
All tests should pass, indicating that the TreeSitterPlugin can successfully instantiate your extractor and parse files with the new language extensions.
Summary
- Create an extractor class under
packages/core/src/plugins/extractors/that implementsLanguageExtractor, importing traversal utilities (traverse,findChild,getStringValue) frombase-extractor.ts. - Register the extractor by exporting it from
packages/core/src/plugins/extractors/index.tsso the plugin system can discover it. - Define language metadata in a new config file under
packages/core/src/languages/configs/specifying the language ID, extensions, and Tree-Sitter grammar dependency, then add it tobuiltinLanguageConfigs. - Install dependencies by adding the Tree-Sitter grammar npm package to
packages/core/package.jsonand runningpnpm install. - Validate with tests by creating a test file in
__tests__/and runningpnpm --filter @understand-anything/core testto ensure the extraction logic works correctly.
Frequently Asked Questions
What is the purpose of base-extractor.ts in Understand-Anything?
The base-extractor.ts module provides shared AST traversal utilities including traverse, getStringValue, findChild, findChildren, and hasChildOfType. These functions encapsulate common Tree-Sitter navigation patterns, allowing language-specific extractors to focus on grammar-specific node handling without duplicating recursion logic or text extraction code.
How does the TreeSitterPlugin discover new extractors?
The TreeSitterPlugin discovers extractors through the central export registry in packages/core/src/plugins/extractors/index.ts. When you export your extractor class from this file, the plugin can instantiate it and match it to files based on the languageIds array and the corresponding entry in builtinLanguageConfigs.
What values should I use for the treeSitter.grammar field in the language config?
The treeSitter.grammar field should contain the npm package name of the Tree-Sitter grammar for your target language (for example, tree-sitter-my-lang). This package must be installed in packages/core/package.json so the parser can load the compiled wasm or native grammar at runtime.
Can I reuse base-extractor helpers for both structure and call-graph extraction?
Yes. The helpers in base-extractor.ts are designed to work with any Tree-Sitter AST. Use traverse for full-tree walks during structural analysis, and findChild or findChildren for targeted node lookups when identifying call expressions. These utilities operate on TreeSitterNode objects regardless of the underlying language grammar.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →