How to Add Support for New Programming Languages to Understand-Anything

To add support for a new programming language in Understand-Anything, create a language configuration in packages/core/src/languages/configs/, register it in the builtin registry, implement a Tree-Sitter extractor extending BaseExtractor, and install the corresponding Tree-Sitter WASM grammar package.

Understand-Anything is an open-source codebase visualization tool that combines Tree-Sitter parsers with language-specific extractors to analyze and map software architecture. When you need to add support for new programming languages to Understand-Anything, you extend the core analysis engine by defining language metadata, registering configurations in the builtin registry, and implementing AST extractors that map language-specific constructs to generic structural models.

Step 1: Define a Language Configuration

Create a new configuration file under packages/core/src/languages/configs/. This file exports a LanguageConfig object that defines the language ID, file extensions, Tree-Sitter WASM module location, and file pattern heuristics.

For example, to add Scala support:

// packages/core/src/languages/configs/scala.ts
import type { LanguageConfig } from "../types.js";

export const scalaConfig = {
  id: "scala",
  displayName: "Scala",
  extensions: [".scala", ".sc"],
  treeSitter: {
    wasmPackage: "tree-sitter-scala",
    wasmFile: "tree-sitter-scala.wasm",
  },
  concepts: [
    "case classes",
    "traits",
    "pattern matching",
  ],
  filePatterns: {
    entryPoints: ["src/main/scala/**/*.scala"],
    tests: ["src/test/scala/**/*.scala"],
    config: ["build.sbt"],
  },
} satisfies LanguageConfig;

Key requirements for the configuration:

  • id must be unique and match the ID used by the extractor class.
  • extensions define which file types trigger this language parser.
  • treeSitter.wasmPackage specifies the npm package containing the grammar; this package must be installed as a dependency in step 4.

Reference: See the existing typescriptConfig in packages/core/src/languages/configs/typescript.ts for the complete configuration schema.

Step 2: Register the Configuration in the Built-In Registry

Edit packages/core/src/languages/configs/index.ts to import your new config and append it to the builtinLanguageConfigs array. This array is the source of truth for which languages the core recognizes.

// packages/core/src/languages/configs/index.ts
import { scalaConfig } from "./scala.js";

export const builtinLanguageConfigs: LanguageConfig[] = [
  // … existing entries …
  scalaConfig,
];

Export the config at the bottom of the file so other modules can import it directly:

export {
  // … existing exports …
  scalaConfig,
};

The LanguageRegistry.createDefault() method initializes the registry using this builtinLanguageConfigs array, automatically recognizing files based on the extensions and patterns you defined.

Step 3: Implement a Tree-Sitter Extractor

Extractors translate language-specific AST nodes into the generic StructureNode and CallGraphNode types used by the analyzer. Create a new extractor under packages/core/src/plugins/extractors/.

// packages/core/src/plugins/extractors/scala-extractor.ts
import {
  BaseExtractor,
  getStringValue,
  findChild,
  findChildren,
} from "./base-extractor.js";
import type { StructureNode, CallGraphNode } from "../../types.js";

export class ScalaExtractor extends BaseExtractor {
  languageIds = ["scala"];

  extractStructure(root: any): StructureNode[] {
    const nodes: StructureNode[] = [];
    const defs = findChildren(root, "class_declaration")
      .concat(findChildren(root, "object_declaration"))
      .concat(findChildren(root, "trait_declaration"));

    for (const def of defs) {
      const name = getStringValue(findChild(def, "type_identifier"));
      nodes.push({
        name,
        kind: "class",
        range: this.nodeRange(def),
        children: [],
      });
    }
    return nodes;
  }

  extractCallGraph(root: any): CallGraphNode[] {
    const calls: CallGraphNode[] = [];
    const invocations = findChildren(root, "method_invocation");

    for (const call of invocations) {
      const callee = getStringValue(findChild(call, "identifier"));
      calls.push({
        caller: this.enclosingFunction(call),
        callee,
        location: this.nodeRange(call),
      });
    }
    return calls;
  }
}

Add the extractor to the exports in packages/core/src/plugins/extractors/index.ts:

export { ScalaExtractor } from "./scala-extractor.js";

The AnalyzerPluginRegistry automatically maps language IDs to extractors when they are included in the default set, enabling automatic analysis without additional configuration.

Step 4: Install the Tree-Sitter Grammar Package

Add the Tree-Sitter grammar package to packages/core/package.json under dependencies:

{
  "dependencies": {
    "tree-sitter-scala": "^0.1.0"
  }
}

Run pnpm install to fetch the WASM file. The TreeSitterPlugin automatically loads the package based on the treeSitter.wasmPackage field defined in your language config.

Step 5: Verify Registration and Test

After completing the previous steps, verify that the LanguageRegistry recognizes your new language:

// packages/core/src/__tests__/language-registry.test.ts
import { LanguageRegistry } from "../languages/language-registry.js";

test("detects Scala files", () => {
  const reg = LanguageRegistry.createDefault();
  const cfg = reg.getForFile("src/main/scala/Hello.scala");
  expect(cfg).not.toBeNull();
  expect(cfg?.id).toBe("scala");
});

Run the core test suite to ensure registration works correctly:

pnpm --filter @understand-anything/core test

The TreeSitterPlugin (located in packages/core/src/plugins/tree-sitter-plugin.ts) automatically instantiates your extractor when it encounters files matching your configured extensions, passing the parsed AST to your extractStructure and extractCallGraph methods.

Optional: Add UI Labels for the Dashboard

If you want the new language to appear in the dashboard's language selector, update the localization files under packages/dashboard/src/locales/. Add a translation entry mapping your language ID to its display name.

Summary

  • Create a language configuration in packages/core/src/languages/configs/{language}.ts defining extensions, Tree-Sitter package, and file patterns.
  • Register the configuration in packages/core/src/languages/configs/index.ts by adding it to the builtinLanguageConfigs array.
  • Implement an extractor extending BaseExtractor in packages/core/src/plugins/extractors/ with extractStructure and extractCallGraph methods.
  • Export the extractor from packages/core/src/plugins/extractors/index.ts so the plugin registry can discover it.
  • Install the Tree-Sitter grammar package via npm/pnpm to provide the WASM parsing module.
  • Verify that LanguageRegistry.createDefault() correctly identifies files and that tests pass.

Frequently Asked Questions

What is the minimum code required to add a new language?

You need three essential components: a LanguageConfig object exported from packages/core/src/languages/configs/{language}.ts, an entry in the builtinLanguageConfigs array in packages/core/src/languages/configs/index.ts, and an extractor class extending BaseExtractor that implements extractStructure and extractCallGraph. Additionally, you must install the corresponding Tree-Sitter grammar package from npm.

Why does my extractor need to extend BaseExtractor?

The BaseExtractor class (located in packages/core/src/plugins/extractors/base-extractor.ts) provides utility methods like nodeRange, enclosingFunction, findChild, and findChildren that simplify AST traversal. It also standardizes the interface expected by TreeSitterPlugin, ensuring your extractor receives the correct root node and context during analysis.

Can I support multiple language versions or dialects?

Yes. You can define multiple configurations with distinct IDs (such as python2 and python3) pointing to different Tree-Sitter grammars, or create a single extractor that handles multiple languageIds by checking node types dynamically within your extractStructure method. The builtinLanguageConfigs array can contain multiple entries that share the same extractor class if their AST structures are compatible.

How do I debug why my language isn't being detected?

First, verify that your config is imported and added to builtinLanguageConfigs in packages/core/src/languages/configs/index.ts. Second, check that the file extension matches the extensions array in your config (extensions are normalized to include a leading dot). Finally, add a unit test calling LanguageRegistry.createDefault().getForFile() with a sample path to confirm the registry resolution logic works as expected.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →