# How to Add Custom Language Support to Code‑Review‑Graph Using Languages.toml

> Easily add custom language support to code-review-graph using languages.toml. Define new languages without modifying source code for enhanced code review analysis.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: how-to-guide
- Published: 2026-08-14

---

**Yes, code‑review‑graph supports custom language definitions through a [`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml) configuration file without requiring any source code modifications.**

The tool automatically discovers and loads user‑defined languages from `.code‑review‑graph/languages.toml` at the repository root. This external configuration lets you register new parsers, file extensions, and language aliases that integrate seamlessly with the analysis pipeline.

## How Languages.toml Discovery Works

The configuration loader in [`code_review_graph/custom_languages.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/custom_languages.py) implements a specific search pattern. When you initialize an analysis, the engine:

1. Checks for a `.code‑review‑graph/` directory at the repository root
2. Looks for a file named exactly [`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml) inside that directory
3. Parses the TOML content and validates each `[languages.<name>]` section
4. Registers valid languages into the parser registry for immediate use

If the file is missing or empty, the analysis proceeds with built‑in languages only. Malformed configurations raise descriptive errors before any code parsing begins.

## Required Schema for Languages.toml

Each language entry under `[languages.<name>]` must specify core fields that the parser requires:

| Field | Type | Required | Purpose |
|-------|------|----------|---------|
| `grammar` | string | Yes | Path to a compiled tree‑sitter WASM grammar or compatible parser module |
| `extensions` | list of strings | Yes | File extensions (without dots) to associate with this language |
| `aliases` | list of strings | No | Alternative identifiers for import‑resolution heuristics |
| `probe_fields` | list of strings | No | AST node attributes to examine when detecting call relationships |

The `grammar` path is resolved relative to the repository root. This design keeps parser assets version‑controlled alongside your configuration.

## Complete Languages.toml Example

Here is a production‑ready configuration adding support for a domain‑specific language named **FlowScript**:

```toml

# File: .code-review-graph/languages.toml

[languages.flowscript]
grammar = "parsers/flowscript/tree-sitter-flowscript.wasm"
extensions = ["flow", "fls"]
aliases = ["flow", "fs"]
probe_fields = ["step", "transition", "handler"]

```

With this file in place, code‑review‑graph will:

- Parse all `.flow` and `.fls` files using the specified WASM grammar
- Recognize `import` statements referencing `flow` or `fs` as FlowScript dependencies
- Extract call relationships from `step`, `transition`, and `handler` AST nodes during graph construction

## Loading Custom Languages Programmatically

The same configuration system is exposed through the Python API defined in [`code_review_graph/custom_languages.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/custom_languages.py). Use this to inspect or debug loaded languages:

```python
from pathlib import Path
from code_review_graph.custom_languages import load_custom_languages

repo_root = Path("/home/user/my-project")
languages = load_custom_languages(repo_root)

for name, config in languages.items():
    print(f"{name}: {config.grammar} handles {config.extensions}")

```

The `load_custom_languages()` function returns a mapping of language identifiers to validated configuration objects. Each object provides typed access to `grammar`, `extensions`, `aliases`, and `probe_fields` attributes.

## Testing Your Custom Language Configuration

The project's test suite in [`tests/test_custom_languages.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_custom_languages.py) demonstrates validation patterns you can adapt. Here is a minimal test ensuring your configuration loads correctly:

```python
from pathlib import Path
from code_review_graph.custom_languages import load_custom_languages

def test_flowscript_configuration(tmp_path: Path):
    # Create mock repository structure

    config_dir = tmp_path / ".code-review-graph"
    config_dir.mkdir()
    
    config_file = config_dir / "languages.toml"
    config_file.write_text("""
[languages.flowscript]
grammar = "parsers/flowscript.wasm"
extensions = ["flow"]
aliases = ["fs"]
probe_fields = ["step"]
""")
    
    # Load and verify

    languages = load_custom_languages(tmp_path)
    assert "flowscript" in languages
    assert languages["flowscript"].extensions == ["flow"]
    assert languages["flowscript"].probe_fields == ["step"]

```

This pattern mirrors the actual validation performed by the tool, helping you catch configuration errors before running full analysis.

## Common Configuration Pitfalls

- **Missing grammar file**: The path in [`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml) must resolve to an existing file; the loader validates filesystem presence
- **Duplicate language names**: Each `[languages.<name>]` table key must be unique across the file
- **Empty extension lists**: At least one extension is required for file discovery to function
- **Directory nesting**: The configuration must reside at `.code‑review‑graph/languages.toml` exactly; alternate paths are not searched

## Summary

- Create `.code‑review‑graph/languages.toml` at your repository root to add custom languages without code changes
- Define each language in a `[languages.<name>]` section with required `grammar` and `extensions` fields
- Optional `aliases` and `probe_fields` refine import resolution and call‑graph accuracy
- The loader in [`code_review_graph/custom_languages.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/custom_languages.py) validates, parses, and registers configurations automatically
- Test your setup using patterns from [`tests/test_custom_languages.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_custom_languages.py) before production use

## Frequently Asked Questions

### What file extensions does code‑review‑graph recognize for custom languages?

You specify extensions explicitly in [`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml) under the `extensions` key for each language. The tool does not infer extensions from grammar files; they must be declared as lowercase strings without leading dots, such as `["ts", "tsx"]` for TypeScript variants.

### Can I override built‑in language definitions using languages.toml?

No, the configuration system only adds new languages. Built‑in parsers for languages like Python, JavaScript, and Go remain fixed. To modify behavior for existing languages, you would need to fork and edit [`code_review_graph/builtin_languages.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/builtin_languages.py), then rebuild the package.

### Where should I place the tree‑sitter grammar WASM files?

Store grammar files anywhere within your repository, then reference them with paths relative to the repository root in [`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml). A common convention is `grammars/<language>/tree-sitter-<language>.wasm`. The loader resolves these paths against the directory containing [`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml)'s parent folder.

### Does code‑review‑graph reload languages.toml during incremental analysis?

No, the configuration is parsed once at analysis startup. If you modify [`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml), restart the analysis process to pick up changes. For long‑running server modes, send a SIGHUP or restart the service to refresh language definitions.