How code-review-graph Builds Its Knowledge Graph from Source Code Using Tree-sitter
The code-review-graph tool constructs a language-agnostic knowledge graph by detecting file languages, lazily loading Tree-sitter grammars, parsing source into ASTs, and walking the syntax tree to extract nodes and edges that represent code entities and their relationships.
The code-review-graph repository provides a framework for transforming raw source code into a queryable knowledge graph. By leveraging Tree-sitter's incremental parsing capabilities, it extracts semantic entities and their relationships from multiple programming languages without requiring language-specific compilers. This article examines how the parser transforms source files into nodes and edges using the Tree-sitter AST.
Language Detection via Extensions and Shebangs
Before parsing begins, CodeParser.detect_language identifies the programming language of each file.
Extension-Based Mapping
The parser consults the EXTENSION_TO_LANGUAGE lookup table defined in code_review_graph/parser.py (lines 13-78). This dictionary maps file extensions to their corresponding language identifiers, enabling instant classification of common file types like .py or .js.
Shebang Fallback for Extensionless Files
For files without extensions, the parser falls back to _detect_language_from_shebang (lines 21-88). This method inspects the shebang line (#!/usr/bin/env python3) to determine the appropriate language, ensuring scripts are correctly identified regardless of naming conventions.
Lazy Loading of Tree-sitter Parsers
Once the language is identified, CodeParser._get_parser manages the lifecycle of Tree-sitter grammar instances.
Per-Process Caching
The method first checks self._parsers, a per-process cache that avoids reloading grammars for files of the same language. This optimization reduces memory overhead and eliminates redundant initialization costs when processing large repositories.
Safe Grammar Loading with Subprocess Probing
If a grammar isn't cached, _load_tree_sitter_parser (lines 60-71) handles the loading. Before importing the native binary, _run_parser_load_probe (lines 75-88) executes a short-lived subprocess to verify the grammar can be loaded without crashing the main process. Upon successful verification, it imports tree_sitter_language_pack.get_parser (line 65) and stores the result in the cache.
AST Parsing and Entity Extraction
With a valid parser instance, the system converts source code into graph elements.
From Bytes to Syntax Tree
The parse_bytes method (lines 99-103 in parser.py) accepts raw file content and produces a Tree-sitter syntax tree:
parser = self._get_parser(language) # Load the TS parser
tree = parser.parse(source) # Generate TS AST
Walking the Tree for Semantic Entities
The private method _walk_tree recursively visits every node in the Tree-sitter AST. It compares each node's type against language-specific tables including _CLASS_TYPES, _FUNCTION_TYPES, _IMPORT_TYPES, and _CALL_TYPES (defined in constants.py and referenced in parser.py lines 98-162).
For each matched node, the walker creates:
- NodeInfo objects representing classes, functions, and types
- EdgeInfo objects representing relationships like CALLS, IMPORTS_FROM, INHERITS, and CONTAINS
Dead Code Detection
During the walk, _is_in_static_dead_guard (lines 88-138) identifies code blocks that are statically unreachable (such as if False: constructs). This prevents the parser from adding spurious CALL edges to functions that are never actually invoked at runtime.
Graph Normalization and Assembly
Path Normalization
Both NodeInfo and EdgeInfo objects undergo path normalization via normalize_file_path (lines 50-66). This converts all file paths to POSIX-style separators, ensuring the knowledge graph remains portable across Windows and Unix-like systems.
Aggregating Results into the Graph
The parse_bytes method returns two Python lists: nodes and edges. These collections are passed to code_review_graph.graph.Graph, which merges per-file results into an in-memory adjacency structure. The Graph class handles deduplication and provides methods for persisting the structure to JSON or SQLite.
Repository-Wide Graph Construction
The top-level build_or_update_graph function in code_review_graph/tools/build.py (line 465) orchestrates the complete pipeline. It iterates over every source file in the repository, invokes CodeParser for each, and accumulates the results into a single Graph instance. Once complete, the graph can be queried via the CLI (code_review_graph/cli.py) or exported for visualization.
Code Examples
Parsing Individual Files
from pathlib import Path
from code_review_graph.parser import CodeParser
# Create a parser for a repo (optional repo_root for custom languages)
parser = CodeParser(repo_root=Path("/my/project"))
# Parse a single file – returns lists of NodeInfo and EdgeInfo objects
file_path = Path("src/example.py")
nodes, edges = parser.parse_file(file_path)
print("Nodes extracted:")
for n in nodes:
print(f"{n.kind} {n.name} ({n.file_path}:{n.line_start})")
print("\nEdges extracted:")
for e in edges:
print(f"{e.kind} {e.source} → {e.target} (line {e.line})")
Building the Complete Graph Programmatically
from code_review_graph.graph import Graph
from code_review_graph.parser import CodeParser
import os
parser = CodeParser()
graph = Graph()
for root, _, files in os.walk("."):
for fn in files:
p = Path(root, fn)
if parser.detect_language(p) is None:
continue
nodes, edges = parser.parse_file(p)
graph.add_nodes(nodes)
graph.add_edges(edges)
graph.dump("graph.json") # Persist for later queries / visualization
Command Line Usage
# Build the entire graph for the current repository
crg build # Equivalent to code_review_graph/tools/build.py
Summary
- Language Detection:
CodeParser.detect_languageuses extension mapping (EXTENSION_TO_LANGUAGE) and shebang inspection (_detect_language_from_shebang) to classify files. - Parser Management:
_get_parserimplements lazy loading with per-process caching and subprocess probing (_run_parser_load_probe) to safely initialize Tree-sitter grammars. - AST Extraction:
parse_bytesgenerates Tree-sitter trees, while_walk_treeextractsNodeInfoandEdgeInfoobjects using language-specific node type tables. - Graph Construction: The
Graphclass aggregates normalized nodes and edges, deduplicating entities across files and supporting export to JSON/SQLite. - Entry Point:
build_or_update_graphintools/build.pycoordinates full-repository scanning and graph assembly.
Frequently Asked Questions
What is the role of Tree-sitter in code-review-graph?
Tree-sitter provides the incremental parsing engine that converts source code into concrete syntax trees. The code-review-graph parser uses these ASTs to identify classes, functions, imports, and calls without requiring language-specific compilers or static analysis tools, enabling support for multiple languages through a unified interface.
How does code-review-graph handle language detection for files without extensions?
For extensionless files, the parser executes _detect_language_from_shebang, which reads the shebang line (e.g., #!/usr/bin/env python3) to determine the interpreter. This allows the system to correctly identify and parse script files that lack standard file extensions.
Why does the parser use a subprocess probe before loading Tree-sitter grammars?
The _run_parser_load_probe method creates a short-lived subprocess to test-load the native Tree-sitter grammar binary. This isolation prevents corrupted or incompatible grammar files from crashing the main parsing process, ensuring robust operation when processing repositories with potentially damaged language packs.
How are file paths normalized in the knowledge graph?
The normalize_file_path function (lines 50-66 in parser.py) converts all file paths to POSIX-style forward slashes during NodeInfo and EdgeInfo creation. This normalization ensures the graph is portable between Windows and Unix-like operating systems, eliminating path separator inconsistencies in the output data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →