What Language Parsers Does code-review-graph Support and How Are They Implemented?
The code-review-graph repository supports approximately 60 programming languages through Tree-Sitter grammars, featuring automatic language detection via file extensions and shebang interpreters, a probing cache system for performant parser loading, and regex-based fallbacks for languages lacking native grammars.
The code-review-graph tool extracts syntactic entities—such as classes, functions, imports, and call sites—from source files to construct code review graphs. To accomplish this across a diverse codebase, the project implements a comprehensive language parser architecture that balances broad language support with extensibility and performance.
Supported Languages and Extension Mapping
The system recognizes languages through a comprehensive mapping table that connects file extensions (and shebang interpreters) to Tree-Sitter language identifiers.
The EXTENSION_TO_LANGUAGE Constant
In code_review_graph/parser.py, the constant EXTENSION_TO_LANGUAGE (lines 713-785) defines the built-in language registry covering approximately 60 languages. The mapping includes:
-
Mainstream languages: Python (
.py), JavaScript/TypeScript (.js,.ts,.tsx), Go (.go), Rust (.rs), Java (.java), C/C++ (.c,.cpp), C# (.cs), PHP (.php), Kotlin (.kt), Swift (.swift), Scala (.scala), Solidity (.sol), Dart (.dart), Lua (.lua), PowerShell (.ps1), Julia (.jl), and Verilog (.v,.sv). -
Configuration and data formats: YAML, TOML, HCL, and Jupyter notebooks (
.ipynb). -
Niche languages: Rescript, VB.NET, GDScript, Nix, Zig, and Objective-C (
.m).
The dictionary also handles shebang-based detection for scripts without standard extensions, ensuring that executable Perl or Ruby scripts are correctly identified even when lacking a .pl or .rb suffix.
How Parsers Are Loaded and Cached
Once a language is identified, the system dynamically loads the appropriate Tree-Sitter parser while avoiding the performance penalty of repeated initialization.
The _load_tree_sitter_parser Function
The core loading logic resides in _load_tree_sitter_parser(grammar) (lines 60-145 of code_review_graph/parser.py). This function implements a robust probing mechanism:
- Process-wide caching: The module maintains
_PARSER_PROBE_RESULTSto store successful parser loads, preventing redundant initialization across multiple file parses. - Subprocess probing: For uncached grammars, the code spawns a probe subprocess that imports
tree_sitter_language_packand requests the specific grammar. This isolation prevents the main process from crashing due to missing native dependencies. - Timeout protection: The probe respects
CRG_PARSER_LOAD_TIMEOUT_SECONDS(defaulting to 5 seconds), ensuring the system remains responsive even when grammars are slow to load or missing. - Failure memoization: Failed attempts (missing native grammars, import errors) are cached to skip costly future probes for unsupported languages.
Upon successful probe completion, the actual parser object is retrieved via language_pack.get_parser(grammar).
Semantic Node-Type Mappings
After obtaining a Tree-Sitter parser, the system must translate raw syntax tree nodes into higher-level semantic concepts used by the graph builder.
Mapping Tree-Sitter Nodes to Graph Entities
In code_review_graph/parser.py (lines 998-1195), the code defines language-specific dictionaries that classify node types:
_CLASS_TYPES: Maps Tree-Sitter node types (e.g.,class_definition,class_declaration) to the graph's class entity concept._FUNCTION_TYPES: Identifies function definitions across languages (e.g.,function_definition,method_declaration,func_literal)._IMPORT_TYPES: Recognizes import and require statements._CALL_TYPES: Detects function call expressions and invocations.
These dictionaries allow the parser to walk the syntax tree and emit standardized NodeInfo and EdgeInfo objects regardless of the source language's specific grammar.
Regex-Based Fallback Parsers
When Tree-Sitter grammars are unavailable for certain languages, the system falls back to hand-rolled regex extractors to maintain basic parsing capability.
Hand-Rolled Extractors for Grammar Gaps
The parser.py module implements specialized extraction functions for languages lacking official Tree-Sitter support:
_parse_rescript: Extracts module and function definitions from Rescript (.res) files using targeted regex patterns._VBNET_*functions: A collection of regex tables and parsers for VB.NET syntax (.vbfiles)._parse_sql: Basic SQL statement extraction for database migration and schema files._extract_hcl_constructs: Handles HashiCorp Configuration Language (Terraform) constructs.
These fallbacks are invoked automatically when the language identifier exists in EXTENSION_TO_LANGUAGE but tree_sitter_language_pack reports no available grammar for that language.
Custom Language Extensions
For specialized workflows, users can extend the built-in language set without modifying core source code.
Adding Languages via languages.toml
The module code_review_graph/custom_languages.py provides the load_custom_languages() function, which reads user-defined languages.toml files. These definitions are merged into the built-in EXTENSION_TO_LANGUAGE mapping, with built-in names taking precedence over custom ones to prevent accidental overrides of standard parsers.
This architecture allows teams to add support for proprietary domain-specific languages or experimental grammars while maintaining compatibility with the core Tree-Sitter infrastructure.
Practical Usage Example
from pathlib import Path
from code_review_graph.parser import CodeReviewGraphParser
# Initialize the parser (loads Tree-Sitter grammars lazily)
crg = CodeReviewGraphParser()
# Parse a Python file – Tree-Sitter "python" grammar is used
py_nodes = crg.parse(Path("example.py"))
# Parse a Go source file – uses the "go" grammar
go_nodes = crg.parse(Path("cmd/main.go"))
# Parse a Jupyter notebook – uses the "notebook" handler
nb_nodes = crg.parse(Path("analysis.ipynb"))
# Parse Rescript (no Tree-Sitter grammar) – falls back to regex
res_nodes = crg.parse(Path("module.res"))
The parser automatically detects the language via EXTENSION_TO_LANGUAGE or shebang interpreter, loads (or reuses) the Tree-Sitter parser through _load_tree_sitter_parser(), and walks the syntax tree using the _CLASS_TYPES, _FUNCTION_TYPES, and related mappings to emit graph nodes and edges.
Summary
- code-review-graph supports approximately 60 languages through the
EXTENSION_TO_LANGUAGEmapping incode_review_graph/parser.py. - Parsers are loaded via
_load_tree_sitter_parser(), which uses subprocess probing with a 5-second timeout and process-wide caching via_PARSER_PROBE_RESULTS. - Tree-Sitter grammars provide the primary parsing mechanism, with node types translated to semantic entities using
_CLASS_TYPES,_FUNCTION_TYPES,_IMPORT_TYPES, and_CALL_TYPESdictionaries. - Regex-based fallbacks in
parser.pyhandle languages like Rescript, VB.NET, SQL, and HCL when Tree-Sitter grammars are unavailable. - Users can extend support through
languages.tomlfiles processed bycode_review_graph/custom_languages.py.
Frequently Asked Questions
How does code-review-graph handle languages without Tree-Sitter grammars?
For languages lacking official Tree-Sitter support, such as Rescript, VB.NET, GDScript, SQL, and HCL, the system falls back to hand-rolled regex-based extractors implemented directly in code_review_graph/parser.py. Functions like _parse_rescript and _extract_hcl_constructs use targeted regular expressions to identify classes, functions, and imports without requiring a full syntax tree.
Can I add support for a custom or proprietary language?
Yes. The repository supports custom language definitions through languages.toml files. The load_custom_languages() function in code_review_graph/custom_languages.py loads these definitions and merges them into the built-in EXTENSION_TO_LANGUAGE table, allowing you to map file extensions to Tree-Sitter grammars or custom handlers while preserving the core parsing infrastructure.
What happens if a Tree-Sitter grammar fails to load?
The _load_tree_sitter_parser() function implements defensive error handling. If a grammar fails to load—whether due to missing native libraries or import errors—the failure is cached in _PARSER_PROBE_RESULTS to prevent repeated attempts. The operation is also bounded by CRG_PARSER_LOAD_TIMEOUT_SECONDS (default 5 seconds), ensuring that missing or broken grammars do not hang the parsing process.
How does the parser determine which language to use for a file?
The parser inspects the file extension against the EXTENSION_TO_LANGUAGE dictionary (lines 713-785 in parser.py). For files without extensions, it examines the shebang line (e.g., #!/usr/bin/env python3). This identifier is then passed to _load_tree_sitter_parser() to retrieve the appropriate Tree-Sitter grammar or trigger a regex fallback if no grammar exists.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →