How to Add Support for a New Programming Language to the Egonex-AI Analyzer

To add support for a new programming language to the Egonex-AI analyzer, you must register the language in the central registry, implement a BaseExtractor-derived class that parses the source code into generic node types, and ensure the extractor is discoverable by the plugin system.

The Egonex-AI analyzer uses a modular plugin architecture located in understand-anything-plugin/packages/core to extract code constructs from supported languages. Adding a new language involves three concrete steps: updating the language registry, implementing an extractor that translates Tree-sitter AST nodes into the analyzer's generic knowledge graph format, and wiring the extractor into the plugin discovery system.

Register the Language in the Language Registry

The first step is to declare your language in src/languages/language-registry.ts. This file exports a languageRegistry object that maps language identifiers to their file extensions and extractor modules.

According to the Egonex-AI source code, each entry must specify:

  • A unique language key (e.g., xyz)
  • The human-readable name
  • An array of file extensions that trigger the language
  • The extractor identifier (filename without the .ts suffix)
// understand-anything-plugin/packages/core/src/languages/language-registry.ts
export const languageRegistry: Record<string, LanguageInfo> = {
  ts: {
    name: "TypeScript",
    extensions: [".ts", ".tsx"],
    extractor: "typescript-extractor",
  },
  // Add your new language entry here
  xyz: {
    name: "MyLang",
    extensions: [".xyz"],
    extractor: "my-lang-extractor",
  },
};

Once registered, the core's discovery module (src/plugins/discovery.ts) automatically routes files with matching extensions to your extractor at runtime.

Implement a Custom Extractor

Extractors reside in src/plugins/extractors/ and must extend the BaseExtractor abstract class defined in base-extractor.ts. The TypeScript extractor (typescript-extractor.ts) serves as the reference implementation.

Your extractor must implement the extract method, which receives file content and path, parses the source using Tree-sitter, and returns an array of typed nodes (FileNode, ClassNode, FunctionNode, etc.) defined in src/types.ts.

Key Implementation Requirements

  • Tree-sitter Grammar: The parser relies on WebAssembly grammars loaded via src/plugins/tree-sitter-plugin.ts. Compile your language's Tree-sitter grammar (e.g., from tree-sitter-mylang) and place the WASM file in src/plugins/tree-sitter/.
  • Node Translation: Walk the Tree-sitter syntax tree and convert nodes into the generic types exported from types.ts.
  • Default Export: Export your class as the default export so the registry can load it via dynamic import().
// understand-anything-plugin/packages/core/src/plugins/extractors/my-lang-extractor.ts
import { BaseExtractor } from "./base-extractor";
import { FileNode, ClassNode, FunctionNode } from "../../types";

export default class MyLangExtractor extends BaseExtractor {
  language = "my-lang";

  async extract(fileContent: string, filePath: string) {
    const tree = await this.parseWithTreeSitter(fileContent, "my-lang");
    const nodes: (FileNode | ClassNode | FunctionNode)[] = [];

    const walk = (node: any) => {
      if (node.type === "class_declaration") {
        nodes.push(this.makeClassNode(node, filePath));
      }
      // Handle functions, imports, and other constructs
      for (const child of node.children) walk(child);
    };
    
    walk(tree.rootNode);
    return nodes;
  }
}

Wire the Extractor into the Plugin System

The plugin architecture loads extractors lazily based on the registry entry. No additional registration code is required beyond the default export in your extractor file.

Ensure your extractor file is named to match the extractor value in the registry (e.g., my-lang-extractor.ts for extractor: "my-lang-extractor"). The discovery system uses this identifier to construct the import path dynamically.

Verify with Unit Tests

Create a test file in src/__tests__/ to verify that the registry correctly discovers and instantiates your extractor.

// understand-anything-plugin/packages/core/src/__tests__/my-lang-extractor.test.ts
import { getExtractor } from "../plugins/registry";

test("my-lang extractor is discoverable", async () => {
  const extractor = await getExtractor("my-lang");
  expect(extractor).toBeDefined();
  expect(extractor.language).toBe("my-lang");
});

Run the core test suite to validate your implementation:

pnpm --filter @understand-anything/core test

Summary

  • Update language-registry.ts: Add a mapping that links file extensions to your extractor identifier.
  • Extend BaseExtractor: Create a file in src/plugins/extractors/ that parses Tree-sitter output and returns typed nodes (FileNode, ClassNode, etc.).
  • Export as default: Ensure your extractor class is the default export for dynamic loading.
  • Include Tree-sitter grammar: Place the compiled WASM grammar in the tree-sitter plugin directory.
  • Write tests: Verify discovery and extraction logic in the __tests__ directory.

Frequently Asked Questions

What role does Tree-sitter play in the Egonex-AI analyzer?

Tree-sitter provides the underlying parsing engine. The Egonex-AI analyzer uses Tree-sitter WebAssembly grammars to generate ASTs, which your extractor then traverses and converts into the generic knowledge-graph format. According to the source code, the parseWithTreeSitter method in base-extractor.ts handles the low-level parsing, allowing extractors to focus on semantic translation.

Do I need to modify the discovery logic to add a language?

No. The discovery module in src/plugins/discovery.ts automatically loads extractors based on the languageRegistry configuration. As long as your entry in language-registry.ts points to a valid extractor file name and that file exports a default class extending BaseExtractor, the system discovers it without additional wiring.

What data structures must an extractor return?

Extractors must return arrays of types defined in src/types.ts, such as FileNode, ClassNode, FunctionNode, and ImportNode. These standardized structures allow the analyzer to build a language-agnostic knowledge graph regardless of the source programming language.

How do I handle languages without existing Tree-sitter grammars?

You must first create or obtain a Tree-sitter grammar for the language. Compile it to WebAssembly and place it in src/plugins/tree-sitter/. Then reference the grammar name in your extractor's parseWithTreeSitter call. The analyzer cannot parse languages without a corresponding Tree-sitter WASM module.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →