How to Use Custom Language Parsers to Extend code-review-graph's Language Support

You can extend code-review-graph to any Tree-sitter-supported language by adding a .code-review-graph/languages.toml file that maps file extensions to grammar names and declares which node types represent functions, classes, imports, and calls—no code changes or forks required.

code-review-graph ships with built-in parsers for dozens of languages, but many repositories use niche or emerging languages. Rather than waiting for upstream support or maintaining a fork, the custom language parser feature lets you teach the tool new grammars through pure configuration. According to the tirth8205/code-review-graph source code, this mechanism integrates seamlessly with the existing extraction pipeline—nodes carry your custom language identifier and edges populate exactly like built-in languages.


How Custom Language Discovery Works

When you instantiate CodeParser with a repo_root argument, it automatically searches for custom language definitions before processing any files.

The Entry Point: load_custom_languages

In code_review_graph/parser.py, lines 2453–2457, the constructor triggers discovery:

if repo_root is not None:
    self._custom_languages = load_custom_languages(
        Path(repo_root),
        builtin_extensions=EXTENSION_TO_LANGUAGE,
        builtin_languages=_builtin_language_names(),
    )

The load_custom_languages function (defined in code_review_graph/custom_languages.py, lines 10–21) looks for .code-review-graph/languages.toml at your repository root, validates each entry, and returns a mapping of {name: CustomLanguage}.

Security and Validation Guardrails

The implementation includes defensive measures to protect build stability:

  • Malformed configs never crash—invalid entries are logged and skipped
  • Built-in languages always win—no configuration can override core language mappings
  • Hard cap of 20 custom languages per repository prevents config abuse
  • File caching by modification time and size eliminates redundant re-parsing

The _validate_entry function (lines 82–104 in custom_languages.py) enforces name format rules, extension syntax, grammar availability in tree_sitter_language_pack, and requires at least one non-empty node-type list.


Merging Custom Languages Into the Parser's Core Maps

If custom languages are found, the parser clones its built-in tables and injects your new mappings. This happens in parser.py, lines 2459–2471:

if self._custom_languages:
    self._extension_map = dict(EXTENSION_TO_LANGUAGE)
    self._class_types   = dict(_CLASS_TYPES)
    self._function_types = dict(_FUNCTION_TYPES)
    self._import_types   = dict(_IMPORT_TYPES)
    self._call_types     = dict(_CALL_TYPES)

    for custom in self._custom_languages.values():
        for ext in custom.extensions:
            self._extension_map[ext] = custom.name
        self._class_types[custom.name]   = list(custom.class_node_types)
        self._function_types[custom.name] = list(custom.function_node_types)
        self._import_types[custom.name]   = list(custom.import_node_types)
        self._call_types[custom.name]     = list(custom.call_node_types)

This injection enables four critical capabilities:

  1. Extension resolution—new file extensions map to your custom language name
  2. Function extraction—the walker recognizes your declared function node types
  3. Class/struct extraction—your class node types populate the graph
  4. Dependency analysis—import and call node types enable cross-file edge generation

All downstream analyses—impact radius calculation, search, community detection, and MCP queries—treat your custom language identically to built-in ones.


Grammar Loading and Tree-Sitter Integration

When processing a file, the parser resolves the grammar through _get_parser (lines 72–78 in parser.py):

custom = self._custom_languages.get(language)
grammar = custom.grammar if custom is not None else language
parser = _load_tree_sitter_parser(grammar)

Key insight: The grammar field in your TOML config specifies which Tree-sitter grammar from tree_sitter_language_pack to use. This means you can leverage any of the 100+ pre-packaged grammars without writing a new parser implementation.

The generic Tree-sitter walker in parser.py then processes your file using the node-type lists you provided. No custom Python code is required—the same extraction logic handles Function, Class, IMPORTS_FROM, and CALLS edges for your new language.


Complete Configuration Example: Adding Erlang Support

Here is a working configuration that enables Erlang parsing using the bundled erlang grammar.

Step 1: Create the Configuration File

Create .code-review-graph/languages.toml in your repository root:


# .code-review-graph/languages.toml

[languages.erlang]
extensions = [".erl"]
grammar = "erlang"                     # Provided by tree_sitter_language_pack

function_node_types = ["function_clause"]
class_node_types    = ["record_decl"]
import_node_types   = ["import_attribute"]
call_node_types     = ["call"]
comment = "Erlang support via bundled grammar"

The extensions array determines which files trigger this language. The grammar field must match a name available in tree_sitter_language_pack. The four node-type arrays declare which Tree-sitter AST nodes represent structural elements in your language.

Step 2: Build Your Project

Run the standard build command:

uv run code-review-graph build

The tool automatically detects your configuration, loads the Erlang grammar, and processes any *.erl files. The resulting graph nodes will have language="erlang".

Step 3: Programmatic Access

Access your custom language data through the same API as built-in languages:

from pathlib import Path
from code_review_graph.parser import CodeParser

# Initialise parser for the current repository (detects custom languages automatically)

parser = CodeParser(repo_root=Path('.'))

# Parse a specific Erlang source file

path = Path('src/math_utils.erl')
nodes, edges = parser.parse_file(path)

# Inspect results

for node in nodes:
    if node.language == 'erlang':
        print(f"Found {node.kind} '{node.name}' in {path}")

Step 4: Query the Generated Graph

After building, query your custom language through the graph interface:

from code_review_graph.graph import Graph

g = Graph.load('graph.db')
erlang_funcs = g.nodes.filter(kind='Function', language='erlang')
print(f"Erlang functions: {len(erlang_funcs)}")

Key Implementation Files

Understanding these source files helps debug custom language issues:

File Purpose
code_review_graph/parser.py Core parser class; loads custom languages at lines 2453–2471, merges them into language tables, and implements _get_parser at lines 72–78
code_review_graph/custom_languages.py Defines CustomLanguage dataclass, implements load_custom_languages, validation logic at lines 82–104, and file caching
docs/CUSTOM_LANGUAGES.md Complete schema reference, worked examples, and troubleshooting guide

Summary

  • Custom language parsers extend code-review-graph through declarative TOML configuration—no Python code or forks required
  • The .code-review-graph/languages.toml file maps extensions to Tree-sitter grammars and declares node types for functions, classes, imports, and calls
  • load_custom_languages in custom_languages.py validates and loads configurations with defensive safeguards against malformed entries
  • The parser merges custom mappings into _extension_map, _class_types, _function_types, _import_types, and _call_types in parser.py lines 2459–2471
  • Grammar resolution in _get_parser (lines 72–78) uses your configured grammar name with tree_sitter_language_pack
  • All graph extraction, analysis, and querying features work identically for custom and built-in languages

Frequently Asked Questions

What Tree-sitter grammars are available for custom languages?

Any grammar bundled with tree-sitter-language-pack is usable—over 100 languages including Erlang, Gleam, Zig, and various domain-specific languages. Check the pack's documentation for the exact grammar name to use in your grammar field.

Can I override a built-in language's parser with custom configuration?

No. The implementation explicitly prioritizes built-in languages during the merge process in parser.py. You cannot break core language support through misconfiguration, and you cannot replace built-in parsers with alternative grammars.

What happens if my languages.toml contains errors?

Invalid entries are logged at warning level and skipped. The build continues with any valid custom languages. Validation in custom_languages.py (lines 82–104) checks name format, extension syntax, grammar existence, and requires at least one populated node-type list.

Is there a performance penalty for using custom languages?

Minimal. The configuration file is cached by modification time and size, so re-parsing only occurs when you edit languages.toml. The actual parsing uses the same optimized Tree-sitter paths as built-in languages. The hard limit of 20 custom languages per repository prevents unbounded growth of internal maps.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →