# What Language Parsers Does code-review-graph Support and How Are They Implemented?

> Explore code-review-graph's language parsers, supporting 60+ languages via Tree-Sitter. Discover auto-detection, efficient caching, and regex fallbacks for seamless code analysis.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: how-to-guide
- Published: 2026-08-15

---

**The `code-review-graph` repository supports approximately 60 programming languages through Tree-Sitter grammars, featuring automatic language detection via file extensions and shebang interpreters, a probing cache system for performant parser loading, and regex-based fallbacks for languages lacking native grammars.**

The `code-review-graph` tool extracts syntactic entities—such as classes, functions, imports, and call sites—from source files to construct code review graphs. To accomplish this across a diverse codebase, the project implements a comprehensive **language parser** architecture that balances broad language support with extensibility and performance.

## Supported Languages and Extension Mapping

The system recognizes languages through a comprehensive mapping table that connects file extensions (and shebang interpreters) to Tree-Sitter language identifiers.

### The EXTENSION_TO_LANGUAGE Constant

In [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py), the constant `EXTENSION_TO_LANGUAGE` (lines 713-785) defines the built-in language registry covering approximately 60 languages. The mapping includes:

- **Mainstream languages**: Python (`.py`), JavaScript/TypeScript (`.js`, `.ts`, `.tsx`), Go (`.go`), Rust (`.rs`), Java (`.java`), C/C++ (`.c`, `.cpp`), C# (`.cs`), PHP (`.php`), Kotlin (`.kt`), Swift (`.swift`), Scala (`.scala`), Solidity (`.sol`), Dart (`.dart`), Lua (`.lua`), PowerShell (`.ps1`), Julia (`.jl`), and Verilog (`.v`, `.sv`).

- **Configuration and data formats**: YAML, TOML, HCL, and Jupyter notebooks (`.ipynb`).
- **Niche languages**: Rescript, VB.NET, GDScript, Nix, Zig, and Objective-C (`.m`).

The dictionary also handles shebang-based detection for scripts without standard extensions, ensuring that executable Perl or Ruby scripts are correctly identified even when lacking a `.pl` or `.rb` suffix.

## How Parsers Are Loaded and Cached

Once a language is identified, the system dynamically loads the appropriate Tree-Sitter parser while avoiding the performance penalty of repeated initialization.

### The _load_tree_sitter_parser Function

The core loading logic resides in `_load_tree_sitter_parser(grammar)` (lines 60-145 of [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py)). This function implements a robust probing mechanism:

1. **Process-wide caching**: The module maintains `_PARSER_PROBE_RESULTS` to store successful parser loads, preventing redundant initialization across multiple file parses.
2. **Subprocess probing**: For uncached grammars, the code spawns a probe subprocess that imports `tree_sitter_language_pack` and requests the specific grammar. This isolation prevents the main process from crashing due to missing native dependencies.
3. **Timeout protection**: The probe respects `CRG_PARSER_LOAD_TIMEOUT_SECONDS` (defaulting to **5 seconds**), ensuring the system remains responsive even when grammars are slow to load or missing.
4. **Failure memoization**: Failed attempts (missing native grammars, import errors) are cached to skip costly future probes for unsupported languages.

Upon successful probe completion, the actual parser object is retrieved via `language_pack.get_parser(grammar)`.

## Semantic Node-Type Mappings

After obtaining a Tree-Sitter parser, the system must translate raw syntax tree nodes into higher-level semantic concepts used by the graph builder.

### Mapping Tree-Sitter Nodes to Graph Entities

In [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py) (lines 998-1195), the code defines language-specific dictionaries that classify node types:

- **`_CLASS_TYPES`**: Maps Tree-Sitter node types (e.g., `class_definition`, `class_declaration`) to the graph's class entity concept.
- **`_FUNCTION_TYPES`**: Identifies function definitions across languages (e.g., `function_definition`, `method_declaration`, `func_literal`).
- **`_IMPORT_TYPES`**: Recognizes import and require statements.
- **`_CALL_TYPES`**: Detects function call expressions and invocations.

These dictionaries allow the parser to walk the syntax tree and emit standardized `NodeInfo` and `EdgeInfo` objects regardless of the source language's specific grammar.

## Regex-Based Fallback Parsers

When Tree-Sitter grammars are unavailable for certain languages, the system falls back to hand-rolled regex extractors to maintain basic parsing capability.

### Hand-Rolled Extractors for Grammar Gaps

The [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py) module implements specialized extraction functions for languages lacking official Tree-Sitter support:

- **`_parse_rescript`**: Extracts module and function definitions from Rescript (`.res`) files using targeted regex patterns.
- **`_VBNET_*` functions**: A collection of regex tables and parsers for VB.NET syntax (`.vb` files).
- **`_parse_sql`**: Basic SQL statement extraction for database migration and schema files.
- **`_extract_hcl_constructs`**: Handles HashiCorp Configuration Language (Terraform) constructs.

These fallbacks are invoked automatically when the language identifier exists in `EXTENSION_TO_LANGUAGE` but `tree_sitter_language_pack` reports no available grammar for that language.

## Custom Language Extensions

For specialized workflows, users can extend the built-in language set without modifying core source code.

### Adding Languages via languages.toml

The module [`code_review_graph/custom_languages.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/custom_languages.py) provides the `load_custom_languages()` function, which reads user-defined [`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml) files. These definitions are merged into the built-in `EXTENSION_TO_LANGUAGE` mapping, with built-in names taking precedence over custom ones to prevent accidental overrides of standard parsers.

This architecture allows teams to add support for proprietary domain-specific languages or experimental grammars while maintaining compatibility with the core Tree-Sitter infrastructure.

## Practical Usage Example

```python
from pathlib import Path
from code_review_graph.parser import CodeReviewGraphParser

# Initialize the parser (loads Tree-Sitter grammars lazily)

crg = CodeReviewGraphParser()

# Parse a Python file – Tree-Sitter "python" grammar is used

py_nodes = crg.parse(Path("example.py"))

# Parse a Go source file – uses the "go" grammar

go_nodes = crg.parse(Path("cmd/main.go"))

# Parse a Jupyter notebook – uses the "notebook" handler

nb_nodes = crg.parse(Path("analysis.ipynb"))

# Parse Rescript (no Tree-Sitter grammar) – falls back to regex

res_nodes = crg.parse(Path("module.res"))

```

The parser automatically detects the language via `EXTENSION_TO_LANGUAGE` or shebang interpreter, loads (or reuses) the Tree-Sitter parser through `_load_tree_sitter_parser()`, and walks the syntax tree using the `_CLASS_TYPES`, `_FUNCTION_TYPES`, and related mappings to emit graph nodes and edges.

## Summary

- **code-review-graph** supports approximately **60 languages** through the `EXTENSION_TO_LANGUAGE` mapping in [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py).
- Parsers are loaded via **`_load_tree_sitter_parser()`**, which uses subprocess probing with a **5-second timeout** and process-wide caching via `_PARSER_PROBE_RESULTS`.
- **Tree-Sitter grammars** provide the primary parsing mechanism, with node types translated to semantic entities using **`_CLASS_TYPES`**, **`_FUNCTION_TYPES`**, **`_IMPORT_TYPES`**, and **`_CALL_TYPES`** dictionaries.
- **Regex-based fallbacks** in [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py) handle languages like Rescript, VB.NET, SQL, and HCL when Tree-Sitter grammars are unavailable.
- Users can extend support through **[`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml)** files processed by [`code_review_graph/custom_languages.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/custom_languages.py).

## Frequently Asked Questions

### How does code-review-graph handle languages without Tree-Sitter grammars?

For languages lacking official Tree-Sitter support, such as Rescript, VB.NET, GDScript, SQL, and HCL, the system falls back to hand-rolled regex-based extractors implemented directly in [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py). Functions like `_parse_rescript` and `_extract_hcl_constructs` use targeted regular expressions to identify classes, functions, and imports without requiring a full syntax tree.

### Can I add support for a custom or proprietary language?

Yes. The repository supports custom language definitions through [`languages.toml`](https://github.com/tirth8205/code-review-graph/blob/main/languages.toml) files. The `load_custom_languages()` function in [`code_review_graph/custom_languages.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/custom_languages.py) loads these definitions and merges them into the built-in `EXTENSION_TO_LANGUAGE` table, allowing you to map file extensions to Tree-Sitter grammars or custom handlers while preserving the core parsing infrastructure.

### What happens if a Tree-Sitter grammar fails to load?

The `_load_tree_sitter_parser()` function implements defensive error handling. If a grammar fails to load—whether due to missing native libraries or import errors—the failure is cached in `_PARSER_PROBE_RESULTS` to prevent repeated attempts. The operation is also bounded by `CRG_PARSER_LOAD_TIMEOUT_SECONDS` (default 5 seconds), ensuring that missing or broken grammars do not hang the parsing process.

### How does the parser determine which language to use for a file?

The parser inspects the file extension against the `EXTENSION_TO_LANGUAGE` dictionary (lines 713-785 in [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py)). For files without extensions, it examines the shebang line (e.g., `#!/usr/bin/env python3`). This identifier is then passed to `_load_tree_sitter_parser()` to retrieve the appropriate Tree-Sitter grammar or trigger a regex fallback if no grammar exists.