How to Add Custom Language Support to Code‑Review‑Graph Using Languages.toml
Yes, code‑review‑graph supports custom language definitions through a languages.toml configuration file without requiring any source code modifications.
The tool automatically discovers and loads user‑defined languages from .code‑review‑graph/languages.toml at the repository root. This external configuration lets you register new parsers, file extensions, and language aliases that integrate seamlessly with the analysis pipeline.
How Languages.toml Discovery Works
The configuration loader in code_review_graph/custom_languages.py implements a specific search pattern. When you initialize an analysis, the engine:
- Checks for a
.code‑review‑graph/directory at the repository root - Looks for a file named exactly
languages.tomlinside that directory - Parses the TOML content and validates each
[languages.<name>]section - Registers valid languages into the parser registry for immediate use
If the file is missing or empty, the analysis proceeds with built‑in languages only. Malformed configurations raise descriptive errors before any code parsing begins.
Required Schema for Languages.toml
Each language entry under [languages.<name>] must specify core fields that the parser requires:
| Field | Type | Required | Purpose |
|---|---|---|---|
grammar |
string | Yes | Path to a compiled tree‑sitter WASM grammar or compatible parser module |
extensions |
list of strings | Yes | File extensions (without dots) to associate with this language |
aliases |
list of strings | No | Alternative identifiers for import‑resolution heuristics |
probe_fields |
list of strings | No | AST node attributes to examine when detecting call relationships |
The grammar path is resolved relative to the repository root. This design keeps parser assets version‑controlled alongside your configuration.
Complete Languages.toml Example
Here is a production‑ready configuration adding support for a domain‑specific language named FlowScript:
# File: .code-review-graph/languages.toml
[languages.flowscript]
grammar = "parsers/flowscript/tree-sitter-flowscript.wasm"
extensions = ["flow", "fls"]
aliases = ["flow", "fs"]
probe_fields = ["step", "transition", "handler"]
With this file in place, code‑review‑graph will:
- Parse all
.flowand.flsfiles using the specified WASM grammar - Recognize
importstatements referencingfloworfsas FlowScript dependencies - Extract call relationships from
step,transition, andhandlerAST nodes during graph construction
Loading Custom Languages Programmatically
The same configuration system is exposed through the Python API defined in code_review_graph/custom_languages.py. Use this to inspect or debug loaded languages:
from pathlib import Path
from code_review_graph.custom_languages import load_custom_languages
repo_root = Path("/home/user/my-project")
languages = load_custom_languages(repo_root)
for name, config in languages.items():
print(f"{name}: {config.grammar} handles {config.extensions}")
The load_custom_languages() function returns a mapping of language identifiers to validated configuration objects. Each object provides typed access to grammar, extensions, aliases, and probe_fields attributes.
Testing Your Custom Language Configuration
The project's test suite in tests/test_custom_languages.py demonstrates validation patterns you can adapt. Here is a minimal test ensuring your configuration loads correctly:
from pathlib import Path
from code_review_graph.custom_languages import load_custom_languages
def test_flowscript_configuration(tmp_path: Path):
# Create mock repository structure
config_dir = tmp_path / ".code-review-graph"
config_dir.mkdir()
config_file = config_dir / "languages.toml"
config_file.write_text("""
[languages.flowscript]
grammar = "parsers/flowscript.wasm"
extensions = ["flow"]
aliases = ["fs"]
probe_fields = ["step"]
""")
# Load and verify
languages = load_custom_languages(tmp_path)
assert "flowscript" in languages
assert languages["flowscript"].extensions == ["flow"]
assert languages["flowscript"].probe_fields == ["step"]
This pattern mirrors the actual validation performed by the tool, helping you catch configuration errors before running full analysis.
Common Configuration Pitfalls
- Missing grammar file: The path in
languages.tomlmust resolve to an existing file; the loader validates filesystem presence - Duplicate language names: Each
[languages.<name>]table key must be unique across the file - Empty extension lists: At least one extension is required for file discovery to function
- Directory nesting: The configuration must reside at
.code‑review‑graph/languages.tomlexactly; alternate paths are not searched
Summary
- Create
.code‑review‑graph/languages.tomlat your repository root to add custom languages without code changes - Define each language in a
[languages.<name>]section with requiredgrammarandextensionsfields - Optional
aliasesandprobe_fieldsrefine import resolution and call‑graph accuracy - The loader in
code_review_graph/custom_languages.pyvalidates, parses, and registers configurations automatically - Test your setup using patterns from
tests/test_custom_languages.pybefore production use
Frequently Asked Questions
What file extensions does code‑review‑graph recognize for custom languages?
You specify extensions explicitly in languages.toml under the extensions key for each language. The tool does not infer extensions from grammar files; they must be declared as lowercase strings without leading dots, such as ["ts", "tsx"] for TypeScript variants.
Can I override built‑in language definitions using languages.toml?
No, the configuration system only adds new languages. Built‑in parsers for languages like Python, JavaScript, and Go remain fixed. To modify behavior for existing languages, you would need to fork and edit code_review_graph/builtin_languages.py, then rebuild the package.
Where should I place the tree‑sitter grammar WASM files?
Store grammar files anywhere within your repository, then reference them with paths relative to the repository root in languages.toml. A common convention is grammars/<language>/tree-sitter-<language>.wasm. The loader resolves these paths against the directory containing languages.toml's parent folder.
Does code‑review‑graph reload languages.toml during incremental analysis?
No, the configuration is parsed once at analysis startup. If you modify languages.toml, restart the analysis process to pick up changes. For long‑running server modes, send a SIGHUP or restart the service to refresh language definitions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →