# How to Add Support for a New Programming Language to the Egonex AI Analyzer

> Learn how to add new programming language support to Egonex AI Analyzer. Implement Tree-Sitter extractors and register languages for enhanced code analysis and plugin discovery.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-20

---

**You add support for a new programming language to the Egonex AI analyzer by registering the language in the central registry, implementing a Tree-Sitter-based extractor that extends `BaseExtractor`, and exporting the module so the plugin discovery system can load it at runtime.**

The Egonex AI analyzer processes code through a plugin-based architecture located in `understand-anything-plugin/packages/core`. Adding a new language requires bridging Tree-Sitter grammar parsing with the analyzer's knowledge graph format by mapping file extensions to an extractor that translates AST nodes into standardized symbol objects.

## Register the Language in the Central Registry

The first step is declaring the language in **[`understand-anything-plugin/packages/core/src/languages/language-registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/languages/language-registry.ts)**. This TypeScript map tells the core which file extensions trigger your language and which extractor module to load.

The registry follows this structure:

```typescript
export const languageRegistry: Record<string, LanguageInfo> = {
  // Existing entry
  ts: {
    name: "TypeScript",
    extensions: [".ts", ".tsx"],
    extractor: "typescript-extractor",
  },
  // Add your new entry here
  mylang: {
    name: "MyLang",
    extensions: [".mylang", ".my"],
    extractor: "my-lang-extractor",
  },
};

```

The `extractor` value must match the filename of your implementation (without the `.ts` suffix) located in `src/plugins/extractors/`. Once registered, the **discovery** module at [`src/plugins/discovery.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/plugins/discovery.ts) automatically loads your extractor when it encounters files matching the specified extensions.

## Implement a Tree-Sitter Powered Extractor

Create a new file in **[`src/plugins/extractors/my-lang-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/plugins/extractors/my-lang-extractor.ts)** that extends `BaseExtractor` from `./base-extractor`. Your implementation must override the `extract` method to parse source code and return standardized nodes defined in [`src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/types.ts) (e.g., `FileNode`, `ClassNode`, `FunctionNode`).

The analyzer relies on **Tree-Sitter** WebAssembly grammars for parsing. Your extractor calls `parseWithTreeSitter` (provided by the base class) to generate an AST, then walks the tree to build the knowledge graph.

Here is the required skeleton:

```typescript
import { BaseExtractor } from "./base-extractor";
import { FileNode, ClassNode, FunctionNode } from "../../types";

export default class MyLangExtractor extends BaseExtractor {
  language = "my-lang";

  async extract(fileContent: string, filePath: string): Promise<(FileNode | ClassNode | FunctionNode)[]> {
    const tree = await this.parseWithTreeSitter(fileContent, "my-lang");
    const nodes: (FileNode | ClassNode | FunctionNode)[] = [];

    const walk = (node: any) => {
      if (node.type === "class_declaration") {
        nodes.push(this.makeClassNode(node, filePath));
      }
      if (node.type === "function_declaration") {
        nodes.push(this.makeFunctionNode(node, filePath));
      }
      // Recursively process children
      for (const child of node.children) {
        walk(child);
      }
    };

    walk(tree.rootNode);
    return nodes;
  }

  private makeClassNode(node: any, filePath: string): ClassNode {
    // Transform tree-sitter node to ClassNode
    return {
      type: "class",
      name: node.childForFieldName("name")?.text,
      filePath,
      // ... additional metadata
    };
  }

  private makeFunctionNode(node: any, filePath: string): FunctionNode {
    // Transform tree-sitter node to FunctionNode
    return {
      type: "function",
      name: node.childForFieldName("name")?.text,
      filePath,
      // ... additional metadata
    };
  }
}

```

Key requirements for the extractor:
- **Extend `BaseExtractor`** to inherit `parseWithTreeSitter` and utility methods
- **Export the class as default** (`export default MyLangExtractor`) to enable dynamic imports by the registry
- **Return typed nodes** that conform to the interfaces in [`src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/types.ts)
- **Use the Tree-Sitter grammar** loaded via [`src/plugins/tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/plugins/tree-sitter-plugin.ts), which manages WebAssembly binaries in `src/plugins/tree-sitter/`

## Integrate the Tree-Sitter Grammar

Before your extractor can parse code, you must provide the Tree-Sitter grammar. Compile the grammar for your target language into WebAssembly and place it in **`src/plugins/tree-sitter/`**. The [`tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/tree-sitter-plugin.ts) module loads these grammars and exposes them to `parseWithTreeSitter` using the language identifier you specified in your extractor class.

If your language does not have an existing Tree-Sitter grammar, you must create one following the Tree-Sitter grammar DSL, compile it to WASM, and reference it in the plugin's grammar loader.

## Wire the Extractor into the Plugin System

The plugin system uses **lazy loading** based on the registry entry. No additional registration code is required beyond the `export default` in your extractor file. However, you should verify that the discovery mechanism correctly resolves your module.

Create a test file at **[`src/__tests__/my-lang-extractor.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/__tests__/my-lang-extractor.test.ts)** to validate the integration:

```typescript
import { getExtractor } from "../plugins/registry";

test("my-lang extractor is discoverable", async () => {
  const extractor = await getExtractor("my-lang");
  expect(extractor).toBeDefined();
  expect(extractor.language).toBe("my-lang");
});

```

Run the core test suite to verify your implementation:

```bash
pnpm --filter @understand-anything/core test

```

## Summary

- **Register the language** in [`src/languages/language-registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/languages/language-registry.ts) by mapping file extensions to an extractor name
- **Implement the extractor** in `src/plugins/extractors/` extending `BaseExtractor` and overriding the `async extract(fileContent, filePath)` method
- **Export the extractor class as default** to enable dynamic loading by the discovery module
- **Include Tree-Sitter grammar** compiled to WebAssembly in `src/plugins/tree-sitter/` for parsing support
- **Return standardized nodes** (`FileNode`, `ClassNode`, `FunctionNode`) defined in [`src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/types.ts) to populate the knowledge graph
- **Test the integration** using the registry's `getExtractor` method to ensure the plugin system recognizes your language

## Frequently Asked Questions

### What file structure is required for a new language extractor?

You must create a new TypeScript file in `understand-anything-plugin/packages/core/src/plugins/extractors/` that exports a default class extending `BaseExtractor`. The filename must match the `extractor` value specified in [`language-registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/language-registry.ts) (e.g., [`my-lang-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/my-lang-extractor.ts) for `"extractor": "my-lang-extractor"`). The class must implement the `extract` method and can import type definitions from [`src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/types.ts).

### Does the Egonex AI analyzer require a custom parser for each language?

No. The analyzer uses **Tree-Sitter** as the universal parsing backend. You need to provide a Tree-Sitter grammar compiled to WebAssembly, but you do not need to implement a parser from scratch. The `BaseExtractor` class provides the `parseWithTreeSitter` method that handles the parsing and returns a Tree-Sitter node tree for your extractor to traverse.

### How does the analyzer discover which extractor to use for a specific file?

The discovery module ([`src/plugins/discovery.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/plugins/discovery.ts)) checks file extensions against the `languageRegistry` map in [`src/languages/language-registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/languages/language-registry.ts). When a match is found, it dynamically imports the corresponding extractor module from `src/plugins/extractors/` using the string identifier provided in the registry entry.

### What node types should my extractor return to populate the knowledge graph?

Your extractor should return objects conforming to the interfaces defined in [`src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/types.ts), such as `FileNode`, `ClassNode`, `FunctionNode`, `VariableNode`, or `ImportNode`. These standardized types allow the analyzer to build a consistent knowledge graph across all supported languages, enabling cross-language analysis and code understanding features.