How to Add Support for a New Programming Language to the Egonex AI Analyzer
You add support for a new programming language to the Egonex AI analyzer by registering the language in the central registry, implementing a Tree-Sitter-based extractor that extends BaseExtractor, and exporting the module so the plugin discovery system can load it at runtime.
The Egonex AI analyzer processes code through a plugin-based architecture located in understand-anything-plugin/packages/core. Adding a new language requires bridging Tree-Sitter grammar parsing with the analyzer's knowledge graph format by mapping file extensions to an extractor that translates AST nodes into standardized symbol objects.
Register the Language in the Central Registry
The first step is declaring the language in understand-anything-plugin/packages/core/src/languages/language-registry.ts. This TypeScript map tells the core which file extensions trigger your language and which extractor module to load.
The registry follows this structure:
export const languageRegistry: Record<string, LanguageInfo> = {
// Existing entry
ts: {
name: "TypeScript",
extensions: [".ts", ".tsx"],
extractor: "typescript-extractor",
},
// Add your new entry here
mylang: {
name: "MyLang",
extensions: [".mylang", ".my"],
extractor: "my-lang-extractor",
},
};
The extractor value must match the filename of your implementation (without the .ts suffix) located in src/plugins/extractors/. Once registered, the discovery module at src/plugins/discovery.ts automatically loads your extractor when it encounters files matching the specified extensions.
Implement a Tree-Sitter Powered Extractor
Create a new file in src/plugins/extractors/my-lang-extractor.ts that extends BaseExtractor from ./base-extractor. Your implementation must override the extract method to parse source code and return standardized nodes defined in src/types.ts (e.g., FileNode, ClassNode, FunctionNode).
The analyzer relies on Tree-Sitter WebAssembly grammars for parsing. Your extractor calls parseWithTreeSitter (provided by the base class) to generate an AST, then walks the tree to build the knowledge graph.
Here is the required skeleton:
import { BaseExtractor } from "./base-extractor";
import { FileNode, ClassNode, FunctionNode } from "../../types";
export default class MyLangExtractor extends BaseExtractor {
language = "my-lang";
async extract(fileContent: string, filePath: string): Promise<(FileNode | ClassNode | FunctionNode)[]> {
const tree = await this.parseWithTreeSitter(fileContent, "my-lang");
const nodes: (FileNode | ClassNode | FunctionNode)[] = [];
const walk = (node: any) => {
if (node.type === "class_declaration") {
nodes.push(this.makeClassNode(node, filePath));
}
if (node.type === "function_declaration") {
nodes.push(this.makeFunctionNode(node, filePath));
}
// Recursively process children
for (const child of node.children) {
walk(child);
}
};
walk(tree.rootNode);
return nodes;
}
private makeClassNode(node: any, filePath: string): ClassNode {
// Transform tree-sitter node to ClassNode
return {
type: "class",
name: node.childForFieldName("name")?.text,
filePath,
// ... additional metadata
};
}
private makeFunctionNode(node: any, filePath: string): FunctionNode {
// Transform tree-sitter node to FunctionNode
return {
type: "function",
name: node.childForFieldName("name")?.text,
filePath,
// ... additional metadata
};
}
}
Key requirements for the extractor:
- Extend
BaseExtractorto inheritparseWithTreeSitterand utility methods - Export the class as default (
export default MyLangExtractor) to enable dynamic imports by the registry - Return typed nodes that conform to the interfaces in
src/types.ts - Use the Tree-Sitter grammar loaded via
src/plugins/tree-sitter-plugin.ts, which manages WebAssembly binaries insrc/plugins/tree-sitter/
Integrate the Tree-Sitter Grammar
Before your extractor can parse code, you must provide the Tree-Sitter grammar. Compile the grammar for your target language into WebAssembly and place it in src/plugins/tree-sitter/. The tree-sitter-plugin.ts module loads these grammars and exposes them to parseWithTreeSitter using the language identifier you specified in your extractor class.
If your language does not have an existing Tree-Sitter grammar, you must create one following the Tree-Sitter grammar DSL, compile it to WASM, and reference it in the plugin's grammar loader.
Wire the Extractor into the Plugin System
The plugin system uses lazy loading based on the registry entry. No additional registration code is required beyond the export default in your extractor file. However, you should verify that the discovery mechanism correctly resolves your module.
Create a test file at src/__tests__/my-lang-extractor.test.ts to validate the integration:
import { getExtractor } from "../plugins/registry";
test("my-lang extractor is discoverable", async () => {
const extractor = await getExtractor("my-lang");
expect(extractor).toBeDefined();
expect(extractor.language).toBe("my-lang");
});
Run the core test suite to verify your implementation:
pnpm --filter @understand-anything/core test
Summary
- Register the language in
src/languages/language-registry.tsby mapping file extensions to an extractor name - Implement the extractor in
src/plugins/extractors/extendingBaseExtractorand overriding theasync extract(fileContent, filePath)method - Export the extractor class as default to enable dynamic loading by the discovery module
- Include Tree-Sitter grammar compiled to WebAssembly in
src/plugins/tree-sitter/for parsing support - Return standardized nodes (
FileNode,ClassNode,FunctionNode) defined insrc/types.tsto populate the knowledge graph - Test the integration using the registry's
getExtractormethod to ensure the plugin system recognizes your language
Frequently Asked Questions
What file structure is required for a new language extractor?
You must create a new TypeScript file in understand-anything-plugin/packages/core/src/plugins/extractors/ that exports a default class extending BaseExtractor. The filename must match the extractor value specified in language-registry.ts (e.g., my-lang-extractor.ts for "extractor": "my-lang-extractor"). The class must implement the extract method and can import type definitions from src/types.ts.
Does the Egonex AI analyzer require a custom parser for each language?
No. The analyzer uses Tree-Sitter as the universal parsing backend. You need to provide a Tree-Sitter grammar compiled to WebAssembly, but you do not need to implement a parser from scratch. The BaseExtractor class provides the parseWithTreeSitter method that handles the parsing and returns a Tree-Sitter node tree for your extractor to traverse.
How does the analyzer discover which extractor to use for a specific file?
The discovery module (src/plugins/discovery.ts) checks file extensions against the languageRegistry map in src/languages/language-registry.ts. When a match is found, it dynamically imports the corresponding extractor module from src/plugins/extractors/ using the string identifier provided in the registry entry.
What node types should my extractor return to populate the knowledge graph?
Your extractor should return objects conforming to the interfaces defined in src/types.ts, such as FileNode, ClassNode, FunctionNode, VariableNode, or ImportNode. These standardized types allow the analyzer to build a consistent knowledge graph across all supported languages, enabling cross-language analysis and code understanding features.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →