Tree-Sitter Programming Language Support in code-review-graph: Complete List and Usage Guide
The Tree-sitter parser in code-review-graph supports 25+ programming languages including Python, JavaScript, TypeScript, Go, Rust, Java, C#, Ruby, C/C++, Kotlin, Swift, PHP, and others with bundled grammars in the tree_sitter_language_pack package.
The code-review-graph project by tirth8205 leverages the Tree-sitter parsing library through the tree_sitter_language_pack package to construct language-aware code graphs for automated code review. Understanding which programming languages are supported—and how the extension-to-language mapping works—is essential for developers integrating this tool into their workflows.
Complete List of Tree-Sitter Supported Languages
The repository defines a comprehensive mapping from file extensions to language identifiers in code_review_graph/parser.py (lines 48-87). Below are all programming languages with native Tree-sitter grammar support:
Web and JavaScript Ecosystem
- JavaScript —
.js,.jsx - TypeScript —
.ts,.tsx,.astro - Vue.js —
.vue(single-file components) - Svelte —
.svelte
Systems and Compiled Languages
- Go —
.go - Rust —
.rs - C / C++ —
.c,.h,.cpp,.cc,.cxx,.hpp,.hh - Zig —
.zig - Swift —
.swift
JVM and Enterprise Languages
- Java —
.java - Kotlin —
.kt - Scala —
.scala
Scripting and Dynamic Languages
- Python —
.py - Ruby —
.rb - PHP —
.php - Dart —
.dart - Elixir —
.ex,.exs
Shell and Data Science
- Bash / Shell —
.sh,.bash,.zsh,.ksh(includes Korn and Zsh variants) - PowerShell —
.ps1,.psm1,.psd1 - R —
.r,.R - Julia —
.jl
Additional Supported Languages
-
C# (CSharp) —
.cs -
Solidity —
.sol -
Objective-C —
.m -
Lua —
.lua -
Luau —
.luau
How Language Detection Works
The EXTENSION_TO_LANGUAGE dictionary in code_review_graph/parser.py serves as the central registry for mapping file extensions to Tree-sitter grammar identifiers.
Detecting Language from File Path
from pathlib import Path
from code_review_graph.parser import EXTENSION_TO_LANGUAGE
def detect_language(file_path: str) -> str | None:
ext = Path(file_path).suffix.lower()
return EXTENSION_TO_LANGUAGE.get(ext)
# Usage examples
print(detect_language("example.py")) # → "python"
print(detect_language("script.rb")) # → "ruby"
print(detect_language("component.tsx")) # → "typescript"
Loading and Using Tree-Sitter Parsers
For parsing source code with Tree-sitter, use the internal _load_tree_sitter_parser function:
from code_review_graph.parser import _load_tree_sitter_parser
def parse_source(grammar: str, source: bytes):
parser = _load_tree_sitter_parser(grammar)
if parser is None:
raise RuntimeError(f"No Tree-sitter parser available for {grammar}")
tree = parser.parse(source)
return tree
# Parse a JavaScript file
with open("app.js", "rb") as f:
tree = parse_source("javascript", f.read())
print(tree.root_node.sexp())
Important Distinction: Tree-Sitter vs. Fallback Parsers
Not all languages in the extension mapping use Tree-sitter grammars. According to the source code analysis:
- Tree-sitter supported: Languages with bundled grammars in
tree_sitter_language_pack(the 25+ languages listed above). - Regex-based fallback: Languages like Visual Basic, ReScript, and SQL lack bundled grammars and fall back to regex-based parsing.
- Custom logic: Jupyter notebooks (
.ipynb) use specialized handling rather than direct Tree-sitter parsing.
Extending Language Support
The repository provides code_review_graph/custom_languages.py for adding custom Tree-sitter languages without forking the repository. This allows teams to incorporate proprietary or experimental grammars into their code review pipeline.
Testing and Validation
Two test files verify Tree-sitter language support:
| Test File | Purpose |
|---|---|
tests/test_parser_load_probe.py |
Verifies Tree-sitter grammars are loadable for each supported language |
tests/test_multilang.py |
Contains language-specific test cases for correct parsing across the full language set |
Summary
- code-review-graph supports 25+ programming languages through native Tree-sitter grammars bundled in
tree_sitter_language_pack. - Language detection relies on the
EXTENSION_TO_LANGUAGEmapping incode_review_graph/parser.py(lines 48-87). - Use
_load_tree_sitter_parser(grammar)to instantiate parsers for supported languages. - Some languages fall back to regex-based parsing; verify Tree-sitter support before relying on structural parsing features.
- Extend support through
custom_languages.pyfor proprietary grammar requirements.
Frequently Asked Questions
What happens if I try to parse a language without Tree-sitter support?
The _load_tree_sitter_parser function returns None for unsupported grammars. Your code should handle this case or rely on the fallback regex-based parsers for languages like SQL or Visual Basic.
Can I add support for a language not in the default list?
Yes. The code_review_graph/custom_languages.py module allows you to register custom Tree-sitter grammars without modifying the core repository. Follow the module's documentation for registering new language bindings.
Does code-review-graph support Jupyter notebooks?
Jupyter notebooks (.ipynb) are handled through custom logic rather than direct Tree-sitter parsing. The tool extracts and processes code cells separately from the JSON notebook structure.
How does the project ensure parser quality across all supported languages?
The tests/test_parser_load_probe.py file programmatically verifies that every listed grammar loads successfully, while tests/test_multilang.py contains comprehensive test cases validating correct AST generation for each supported language.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →