What Programming Languages Does Code-Review-Graph Support? (And How to Add More)
Code-review-graph supports any language with a Tree-sitter grammar through built-in mappings for Python, JavaScript, TypeScript, Java, Go, Rust, and 20+ others, plus extensible custom language definitions via YAML configuration.
The code-review-graph tool is a code analysis engine that automatically detects and parses multiple programming languages using Tree-sitter grammars. Understanding what programming languages code-review-graph supports out of the box—and how to extend that support—is essential for teams analyzing polyglot codebases.
Built-In Language Catalog
The core language detection system resides in code_review_graph/parser.py and ships with comprehensive mappings for popular languages.
Extension-to-Language Mapping
The EXTENSION_TO_LANGUAGE dictionary at line 713 defines the canonical mappings from file extensions to language identifiers. This includes support for:
- Python (
.py) - JavaScript (
.js,.mjs) - TypeScript (
.ts) - TSX (
.tsx) - Java (
.java) - Go (
.go) - Rust (
.rs) - C (
.c,.h) - C++ (
.cpp,.cc,.hpp) - PHP (
.php) - C# (
.cs) - Ruby (
.rb) - Julia (
.jl)
The dictionary contains thousands of extension mappings that the parser consults when indexing a repository.
Shebang Detection
For files without extensions, the parser falls back to shebang analysis. The SHEBANG_INTERPRETER_TO_LANGUAGE mapping at line 819 in parser.py interprets hashbang lines like #!/usr/bin/python3 to determine the correct language identifier.
Language Families and Special Handling
Certain languages share analysis logic through family groupings defined in code_review_graph/graph.py (lines 44-46). For example, JavaScript, TypeScript, and TSX are treated as a related family for certain graph operations.
Some analysis features, such as receiver-scope and typed-call resolution, only apply to specific languages. These are listed in code_review_graph/scoped_resolver.py (lines 72-76) where _SCOPED_LANGUAGES = ("php", "rust", "csharp") defines which languages receive specialized scope resolution.
How Language Detection Works
The parser uses a two-phase detection strategy that combines static mappings with dynamic Tree-sitter grammar loading.
When the Parser class initializes, it probes for available Tree-sitter grammars using _parser_load_probe_succeeds and _load_tree_sitter_parser (lines 610-665 in parser.py). If a grammar cannot be loaded for a detected language, the system emits a warning and silently ignores that language, preventing analysis crashes while continuing to process supported files.
The language identifier attached to a file determines how the system generates AST nodes (functions, classes, imports) and aggregates statistics via get_stats().
Adding Custom Language Support
You can extend code-review-graph to support new languages without modifying the core source code.
Custom Language Configuration
Create a custom_languages.yaml (or JSON) file in your repository root to declare additional extensions or override existing mappings. The system enforces a maximum of MAX_CUSTOM_LANGUAGES = 20 custom definitions per repository via code_review_graph/custom_languages.py.
The load_custom_languages() function reads your configuration and returns CustomLanguage objects that the parser merges with built-in mappings. During Parser.__init__ (lines 2455-2459 in parser.py), the system copies the built-in EXTENSION_TO_LANGUAGE dictionary and updates it with your custom definitions.
Tree-Sitter Grammar Requirements
For custom language support to function, a Tree-sitter grammar must be available in the runtime environment. The parser automatically attempts to load the grammar specified in your configuration. If the grammar is present, the language is fully integrated into the graph generation pipeline, including node creation and statistics aggregation.
Working with Languages in Code
Here are practical examples for interacting with language detection programmatically.
Listing Detected Languages
from code_review_graph import Graph
graph = Graph(root_path="path/to/repo")
stats = graph.get_stats()
print("Languages detected:", sorted(stats.languages))
# Output: Languages detected: ['go', 'javascript', 'python', 'rust']
Adding a Custom Language Definition
Create a custom_languages.yaml file:
- name: "mydsl"
extensions: [".mydsl"]
parser: "tree-sitter-mydsl"
Then load the graph with the custom configuration:
from code_review_graph import Graph
graph = Graph(
root_path="path/to/repo",
custom_lang_path="custom_languages.yaml"
)
print(graph.get_stats().languages) # Now includes 'mydsl'
Querying Specific Language Nodes
from code_review_graph import Query
q = Query(graph)
python_funcs = q.nodes.filter(language="python", kind="function")
for fn in python_funcs:
print(f"{fn.file_path}: {fn.name}")
Summary
-
Code-review-graph supports 20+ languages out of the box via
EXTENSION_TO_LANGUAGEinparser.py, including Python, JavaScript, Go, Rust, and Java. -
Language detection uses file extensions, shebang patterns, and Tree-sitter grammar availability, with graceful degradation when grammars are missing.
-
Scoped analysis applies only to specific languages like PHP, Rust, and C# as defined in
scoped_resolver.py. -
You can add custom languages via YAML configuration (up to 20 per repo) without touching core code, provided a Tree-sitter grammar exists.
Frequently Asked Questions
How many programming languages does code-review-graph support by default?
The tool ships with built-in mappings for approximately two dozen major languages, including Python, JavaScript/TypeScript, Java, Go, Rust, C/C++, PHP, C#, Ruby, and Julia. The actual number of parseable languages depends on which Tree-sitter grammars are installed in your environment, as the system can theoretically support any language with an available grammar.
Can I use code-review-graph for a language not in the built-in list?
Yes. You can define custom languages by creating a custom_languages.yaml file that maps file extensions to Tree-sitter parser names. As long as the Tree-sitter grammar is installed in your Python environment, the parser will automatically load and analyze files in that language.
What happens if a Tree-sitter grammar is missing for a detected language?
The parser probes for grammar availability during initialization using _parser_load_probe_succeeds. If a grammar cannot be loaded, the system emits a warning and excludes that language from analysis, but continues processing other files. This prevents crashes while providing visibility into missing dependencies.
Are all analysis features available for every supported language?
No. Certain advanced features like receiver-scope resolution and typed-call analysis are limited to specific languages defined in scoped_resolver.py (currently PHP, Rust, and C#). Basic node extraction and graph building work for all languages with Tree-sitter grammars, but language-specific semantic analysis depends on dedicated implementations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →