How to Add Support for a New Programming Language to the Egonex AI Analyzer

You add support for a new programming language to the Egonex AI analyzer by registering the language in the central registry, implementing a Tree-Sitter-based extractor that extends BaseExtractor, and exporting the module so the plugin discovery system can load it at runtime.

The Egonex AI analyzer processes code through a plugin-based architecture located in understand-anything-plugin/packages/core. Adding a new language requires bridging Tree-Sitter grammar parsing with the analyzer's knowledge graph format by mapping file extensions to an extractor that translates AST nodes into standardized symbol objects.

Register the Language in the Central Registry

The first step is declaring the language in understand-anything-plugin/packages/core/src/languages/language-registry.ts. This TypeScript map tells the core which file extensions trigger your language and which extractor module to load.

The registry follows this structure:

export const languageRegistry: Record<string, LanguageInfo> = {
  // Existing entry
  ts: {
    name: "TypeScript",
    extensions: [".ts", ".tsx"],
    extractor: "typescript-extractor",
  },
  // Add your new entry here
  mylang: {
    name: "MyLang",
    extensions: [".mylang", ".my"],
    extractor: "my-lang-extractor",
  },
};

The extractor value must match the filename of your implementation (without the .ts suffix) located in src/plugins/extractors/. Once registered, the discovery module at src/plugins/discovery.ts automatically loads your extractor when it encounters files matching the specified extensions.

Implement a Tree-Sitter Powered Extractor

Create a new file in src/plugins/extractors/my-lang-extractor.ts that extends BaseExtractor from ./base-extractor. Your implementation must override the extract method to parse source code and return standardized nodes defined in src/types.ts (e.g., FileNode, ClassNode, FunctionNode).

The analyzer relies on Tree-Sitter WebAssembly grammars for parsing. Your extractor calls parseWithTreeSitter (provided by the base class) to generate an AST, then walks the tree to build the knowledge graph.

Here is the required skeleton:

import { BaseExtractor } from "./base-extractor";
import { FileNode, ClassNode, FunctionNode } from "../../types";

export default class MyLangExtractor extends BaseExtractor {
  language = "my-lang";

  async extract(fileContent: string, filePath: string): Promise<(FileNode | ClassNode | FunctionNode)[]> {
    const tree = await this.parseWithTreeSitter(fileContent, "my-lang");
    const nodes: (FileNode | ClassNode | FunctionNode)[] = [];

    const walk = (node: any) => {
      if (node.type === "class_declaration") {
        nodes.push(this.makeClassNode(node, filePath));
      }
      if (node.type === "function_declaration") {
        nodes.push(this.makeFunctionNode(node, filePath));
      }
      // Recursively process children
      for (const child of node.children) {
        walk(child);
      }
    };

    walk(tree.rootNode);
    return nodes;
  }

  private makeClassNode(node: any, filePath: string): ClassNode {
    // Transform tree-sitter node to ClassNode
    return {
      type: "class",
      name: node.childForFieldName("name")?.text,
      filePath,
      // ... additional metadata
    };
  }

  private makeFunctionNode(node: any, filePath: string): FunctionNode {
    // Transform tree-sitter node to FunctionNode
    return {
      type: "function",
      name: node.childForFieldName("name")?.text,
      filePath,
      // ... additional metadata
    };
  }
}

Key requirements for the extractor:

  • Extend BaseExtractor to inherit parseWithTreeSitter and utility methods
  • Export the class as default (export default MyLangExtractor) to enable dynamic imports by the registry
  • Return typed nodes that conform to the interfaces in src/types.ts
  • Use the Tree-Sitter grammar loaded via src/plugins/tree-sitter-plugin.ts, which manages WebAssembly binaries in src/plugins/tree-sitter/

Integrate the Tree-Sitter Grammar

Before your extractor can parse code, you must provide the Tree-Sitter grammar. Compile the grammar for your target language into WebAssembly and place it in src/plugins/tree-sitter/. The tree-sitter-plugin.ts module loads these grammars and exposes them to parseWithTreeSitter using the language identifier you specified in your extractor class.

If your language does not have an existing Tree-Sitter grammar, you must create one following the Tree-Sitter grammar DSL, compile it to WASM, and reference it in the plugin's grammar loader.

Wire the Extractor into the Plugin System

The plugin system uses lazy loading based on the registry entry. No additional registration code is required beyond the export default in your extractor file. However, you should verify that the discovery mechanism correctly resolves your module.

Create a test file at src/__tests__/my-lang-extractor.test.ts to validate the integration:

import { getExtractor } from "../plugins/registry";

test("my-lang extractor is discoverable", async () => {
  const extractor = await getExtractor("my-lang");
  expect(extractor).toBeDefined();
  expect(extractor.language).toBe("my-lang");
});

Run the core test suite to verify your implementation:

pnpm --filter @understand-anything/core test

Summary

  • Register the language in src/languages/language-registry.ts by mapping file extensions to an extractor name
  • Implement the extractor in src/plugins/extractors/ extending BaseExtractor and overriding the async extract(fileContent, filePath) method
  • Export the extractor class as default to enable dynamic loading by the discovery module
  • Include Tree-Sitter grammar compiled to WebAssembly in src/plugins/tree-sitter/ for parsing support
  • Return standardized nodes (FileNode, ClassNode, FunctionNode) defined in src/types.ts to populate the knowledge graph
  • Test the integration using the registry's getExtractor method to ensure the plugin system recognizes your language

Frequently Asked Questions

What file structure is required for a new language extractor?

You must create a new TypeScript file in understand-anything-plugin/packages/core/src/plugins/extractors/ that exports a default class extending BaseExtractor. The filename must match the extractor value specified in language-registry.ts (e.g., my-lang-extractor.ts for "extractor": "my-lang-extractor"). The class must implement the extract method and can import type definitions from src/types.ts.

Does the Egonex AI analyzer require a custom parser for each language?

No. The analyzer uses Tree-Sitter as the universal parsing backend. You need to provide a Tree-Sitter grammar compiled to WebAssembly, but you do not need to implement a parser from scratch. The BaseExtractor class provides the parseWithTreeSitter method that handles the parsing and returns a Tree-Sitter node tree for your extractor to traverse.

How does the analyzer discover which extractor to use for a specific file?

The discovery module (src/plugins/discovery.ts) checks file extensions against the languageRegistry map in src/languages/language-registry.ts. When a match is found, it dynamically imports the corresponding extractor module from src/plugins/extractors/ using the string identifier provided in the registry entry.

What node types should my extractor return to populate the knowledge graph?

Your extractor should return objects conforming to the interfaces defined in src/types.ts, such as FileNode, ClassNode, FunctionNode, VariableNode, or ImportNode. These standardized types allow the analyzer to build a consistent knowledge graph across all supported languages, enabling cross-language analysis and code understanding features.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →