# How to Add Support for a New Programming Language to the Egonex-AI Analyzer

> Learn how to add support for a new programming language to the Egonex-AI analyzer. Register your language, implement a BaseExtractor class, and make it discoverable.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-14

---

**To add support for a new programming language to the Egonex-AI analyzer, you must register the language in the central registry, implement a `BaseExtractor`-derived class that parses the source code into generic node types, and ensure the extractor is discoverable by the plugin system.**

The Egonex-AI analyzer uses a modular plugin architecture located in `understand-anything-plugin/packages/core` to extract code constructs from supported languages. Adding a new language involves three concrete steps: updating the language registry, implementing an extractor that translates Tree-sitter AST nodes into the analyzer's generic knowledge graph format, and wiring the extractor into the plugin discovery system.

## Register the Language in the Language Registry

The first step is to declare your language in **[`src/languages/language-registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/languages/language-registry.ts)**. This file exports a `languageRegistry` object that maps language identifiers to their file extensions and extractor modules.

According to the Egonex-AI source code, each entry must specify:
- A unique language key (e.g., `xyz`)
- The human-readable name
- An array of file extensions that trigger the language
- The extractor identifier (filename without the `.ts` suffix)

```typescript
// understand-anything-plugin/packages/core/src/languages/language-registry.ts
export const languageRegistry: Record<string, LanguageInfo> = {
  ts: {
    name: "TypeScript",
    extensions: [".ts", ".tsx"],
    extractor: "typescript-extractor",
  },
  // Add your new language entry here
  xyz: {
    name: "MyLang",
    extensions: [".xyz"],
    extractor: "my-lang-extractor",
  },
};

```

Once registered, the core's **discovery module** ([`src/plugins/discovery.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/plugins/discovery.ts)) automatically routes files with matching extensions to your extractor at runtime.

## Implement a Custom Extractor

Extractors reside in **`src/plugins/extractors/`** and must extend the `BaseExtractor` abstract class defined in [`base-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/base-extractor.ts). The TypeScript extractor ([`typescript-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/typescript-extractor.ts)) serves as the reference implementation.

Your extractor must implement the `extract` method, which receives file content and path, parses the source using Tree-sitter, and returns an array of typed nodes (`FileNode`, `ClassNode`, `FunctionNode`, etc.) defined in **[`src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/types.ts)**.

### Key Implementation Requirements

- **Tree-sitter Grammar**: The parser relies on WebAssembly grammars loaded via [`src/plugins/tree-sitter-plugin.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/plugins/tree-sitter-plugin.ts). Compile your language's Tree-sitter grammar (e.g., from `tree-sitter-mylang`) and place the WASM file in `src/plugins/tree-sitter/`.
- **Node Translation**: Walk the Tree-sitter syntax tree and convert nodes into the generic types exported from [`types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/types.ts).
- **Default Export**: Export your class as the default export so the registry can load it via dynamic `import()`.

```typescript
// understand-anything-plugin/packages/core/src/plugins/extractors/my-lang-extractor.ts
import { BaseExtractor } from "./base-extractor";
import { FileNode, ClassNode, FunctionNode } from "../../types";

export default class MyLangExtractor extends BaseExtractor {
  language = "my-lang";

  async extract(fileContent: string, filePath: string) {
    const tree = await this.parseWithTreeSitter(fileContent, "my-lang");
    const nodes: (FileNode | ClassNode | FunctionNode)[] = [];

    const walk = (node: any) => {
      if (node.type === "class_declaration") {
        nodes.push(this.makeClassNode(node, filePath));
      }
      // Handle functions, imports, and other constructs
      for (const child of node.children) walk(child);
    };
    
    walk(tree.rootNode);
    return nodes;
  }
}

```

## Wire the Extractor into the Plugin System

The plugin architecture loads extractors lazily based on the registry entry. No additional registration code is required beyond the default export in your extractor file.

Ensure your extractor file is named to match the `extractor` value in the registry (e.g., [`my-lang-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/my-lang-extractor.ts) for `extractor: "my-lang-extractor"`). The discovery system uses this identifier to construct the import path dynamically.

## Verify with Unit Tests

Create a test file in **`src/__tests__/`** to verify that the registry correctly discovers and instantiates your extractor.

```typescript
// understand-anything-plugin/packages/core/src/__tests__/my-lang-extractor.test.ts
import { getExtractor } from "../plugins/registry";

test("my-lang extractor is discoverable", async () => {
  const extractor = await getExtractor("my-lang");
  expect(extractor).toBeDefined();
  expect(extractor.language).toBe("my-lang");
});

```

Run the core test suite to validate your implementation:

```bash
pnpm --filter @understand-anything/core test

```

## Summary

- **Update [`language-registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/language-registry.ts)**: Add a mapping that links file extensions to your extractor identifier.
- **Extend `BaseExtractor`**: Create a file in `src/plugins/extractors/` that parses Tree-sitter output and returns typed nodes (`FileNode`, `ClassNode`, etc.).
- **Export as default**: Ensure your extractor class is the default export for dynamic loading.
- **Include Tree-sitter grammar**: Place the compiled WASM grammar in the tree-sitter plugin directory.
- **Write tests**: Verify discovery and extraction logic in the `__tests__` directory.

## Frequently Asked Questions

### What role does Tree-sitter play in the Egonex-AI analyzer?

Tree-sitter provides the underlying parsing engine. The Egonex-AI analyzer uses Tree-sitter WebAssembly grammars to generate ASTs, which your extractor then traverses and converts into the generic knowledge-graph format. According to the source code, the `parseWithTreeSitter` method in [`base-extractor.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/base-extractor.ts) handles the low-level parsing, allowing extractors to focus on semantic translation.

### Do I need to modify the discovery logic to add a language?

No. The discovery module in [`src/plugins/discovery.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/plugins/discovery.ts) automatically loads extractors based on the `languageRegistry` configuration. As long as your entry in [`language-registry.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/language-registry.ts) points to a valid extractor file name and that file exports a default class extending `BaseExtractor`, the system discovers it without additional wiring.

### What data structures must an extractor return?

Extractors must return arrays of types defined in [`src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/types.ts), such as `FileNode`, `ClassNode`, `FunctionNode`, and `ImportNode`. These standardized structures allow the analyzer to build a language-agnostic knowledge graph regardless of the source programming language.

### How do I handle languages without existing Tree-sitter grammars?

You must first create or obtain a Tree-sitter grammar for the language. Compile it to WebAssembly and place it in `src/plugins/tree-sitter/`. Then reference the grammar name in your extractor's `parseWithTreeSitter` call. The analyzer cannot parse languages without a corresponding Tree-sitter WASM module.