How to Add Support for a Custom Programming Language to the Understand-Anything Analyzer

To add a custom programming language to the Lum1104/Understand-Anything analyzer, create a validated LanguageConfig describing file extensions and concepts, register it with the LanguageRegistry either statically or dynamically, and optionally implement a LanguageExtractor for Tree-Sitter-based structural analysis.

The Understand-Anything analyzer provides a pluggable architecture for analyzing codebases across multiple languages. Adding support for a custom programming language requires defining a LanguageConfig and registering it with the LanguageRegistry, allowing the core PluginRegistry to route files to the appropriate analysis pipeline.

Understanding the Language Discovery Architecture

The analyzer resolves file types through a chain of registries. When processing a file, the PluginRegistry queries LanguageRegistry.getForFile() to identify the language, then selects a compatible AnalyzerPlugin such as the TreeSitterPlugin. Each language is defined by a LanguageConfig object validated against LanguageConfigSchema (located in understand-anything-plugin/packages/core/src/languages/types.ts). Registrations can occur statically through the built-in configuration list or dynamically at runtime.

Step 1: Define the Language Configuration

Create a TypeScript file in understand-anything-plugin/packages/core/src/languages/configs/ that exports a parsed configuration object describing your language's metadata and file patterns.

Creating the Configuration File

The configuration must specify a unique identifier, display name, file extensions, and pattern groups for entry points and tests. Include an optional treeSitter block if you plan to provide a WebAssembly grammar.

// src/languages/configs/my-lang.ts
import { LanguageConfigSchema } from "../types.js";

export const myLangConfig = LanguageConfigSchema.parse({
  id: "myLang",                     // unique identifier
  displayName: "MyLang",           // human-readable name
  extensions: [".my"],             // file extensions (".my" or "my")
  // filenames: ["MyLangFile"]    // optional explicit filename matches
  concepts: ["functions", "modules", "imports"], // tags used by the LLM
  filePatterns: {
    entryPoints: ["src/**/*.my"],
    barrels:      ["src/**/*.my"],          // where symbols are re-exported
    tests:        ["test/**/*.my"],
    config:       ["my-lang.config.json"], // optional config files
  },
  // Optional Tree-Sitter definition – see step 3 if you have a grammar
  // treeSitter: {
  //   wasmPackage: "tree-sitter-mylang",
  //   wasmFile:    "tree-sitter-mylang.wasm",
  // },
});

The schema validation is performed by LanguageConfigSchema to ensure type safety and required field presence.

Step 2: Register the Language with LanguageRegistry

You can integrate your language either by modifying the built-in registry or by providing a custom instance at runtime.

Static Registration (Built-in)

Modify src/languages/configs/index.ts to import and export your configuration within the builtinLanguageConfigs array. This ensures LanguageRegistry.createDefault() automatically includes your language (as implemented in understand-anything-plugin/packages/core/src/languages/language-registry.ts at line 55).

// src/languages/configs/index.ts
import { myLangConfig } from "./my-lang.js";   // ← add this line

export const builtinLanguageConfigs: LanguageConfig[] = [
  // …existing configs…
  myLangConfig,                               // ← add here
];

Dynamic Registration (Runtime)

For plugin-style architectures where modifying the core repository is undesirable, instantiate a registry and call the register() method:

import { LanguageRegistry } from "./languages/language-registry.js";
import { myLangConfig } from "./languages/configs/my-lang.js";

const registry = LanguageRegistry.createDefault(); // loads built-ins
registry.register(myLangConfig);                  // adds your language

Pass this registry to the PluginRegistry constructor (defined at line 16 of understand-anything-plugin/packages/core/src/plugins/registry.ts) so the analysis pipeline uses your custom registration.

Step 3: Enable Deep Structural Analysis with Tree-Sitter (Optional)

To enable parsing beyond simple file identification, provide a Tree-Sitter WebAssembly grammar and implement a LanguageExtractor. This allows the TreeSitterPlugin to analyze syntax trees and extract functions, classes, and imports.

Providing the Grammar and Extractor

Create an extractor implementing the LanguageExtractor interface with methods for extractStructure() and extractCallGraph(), processing TreeNode objects from the web-tree-sitter library:

// src/plugins/extractors/my-lang-extractor.ts
import type { LanguageExtractor } from "./types.js";
import type { TreeNode } from "web-tree-sitter";

export const myLangExtractor: LanguageExtractor = {
  languageIds: ["myLang"],               // language IDs this extractor handles
  extractStructure(root: TreeNode) {     // called by TreeSitterPlugin.analyzeFile
    // Walk the tree and return a StructuralAnalysis object.
    // For a simple example we just return empty arrays:
    return { functions: [], classes: [], imports: [], exports: [] };
  },
  extractCallGraph(root: TreeNode) {
    // Return CallGraphEntry[] if you want call-graph support.
    return [];
  },
};

Initialize the TreeSitterPlugin with your configuration and extractor:

import { TreeSitterPlugin } from "./plugins/tree-sitter-plugin.js";
import { myLangConfig } from "./languages/configs/my-lang.js";
import { myLangExtractor } from "./plugins/extractors/my-lang-extractor.js";

const tsPlugin = new TreeSitterPlugin(
  [myLangConfig],               // language configs that have a treeSitter field
  [myLangExtractor]            // optional custom extractor(s)
);
await tsPlugin.init();          // must be awaited before use

Now any file with the .my extension will be parsed by Tree-Sitter and processed by your extractor.

Step 4: Verify the Integration

Validate that the registry correctly identifies your language's files using the registry's lookup methods:

import { LanguageRegistry } from "./languages/language-registry.js";
import { myLangConfig } from "./languages/configs/my-lang.js";

const reg = LanguageRegistry.createDefault();
reg.register(myLangConfig);

console.assert(reg.getByExtension(".my")?.id === "myLang");
console.assert(reg.getForFile("src/example.my")?.id === "myLang");

The PluginRegistry.getPluginForFile() method (lines 44-46 of understand-anything-plugin/packages/core/src/plugins/registry.ts) will now route matching files to the appropriate analyzer plugin.

Summary

  • Define a LanguageConfig using LanguageConfigSchema.parse() to specify extensions, concepts, and file patterns for your custom language.
  • Register the configuration statically by adding to builtinLanguageConfigs in src/languages/configs/index.ts, or dynamically via LanguageRegistry.register().
  • Optionally implement a LanguageExtractor and provide Tree-Sitter WASM configuration in the treeSitter field for deep structural analysis and call graph extraction.
  • Pass the custom LanguageRegistry to the PluginRegistry constructor to activate the integration throughout the analysis pipeline.

Frequently Asked Questions

What is the minimum configuration required to add a language?

You must provide a unique id, displayName, and at least one file extension in the LanguageConfig. The concepts and filePatterns fields help the LLM context engine but are optional for basic file recognition and routing.

Can I add language support without modifying the core repository?

Yes. You can create a LanguageRegistry instance via LanguageRegistry.createDefault(), call register() with your configuration, and pass this registry to the PluginRegistry constructor. This pattern supports external plugins without editing the built-in configuration files in the Lum1104/Understand-Anything repository.

How does the analyzer match files to my custom language?

The LanguageRegistry uses getForFile() to match against the extensions array and optional filenames list defined in your LanguageConfig. The first matching configuration is returned to the PluginRegistry, which then selects the appropriate AnalyzerPlugin based on the language's capabilities.

Is Tree-Sitter integration mandatory for custom language support?

No. Tree-Sitter integration is only required if you need structural analysis such as function extraction or call graph generation. For simple language identification and LLM-based analysis without AST parsing, the LanguageConfig alone is sufficient for the analyzer to recognize and process your custom language files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →