How code-review-graph Builds Its Knowledge Graph from Source Code Using Tree-sitter

The code-review-graph tool constructs a language-agnostic knowledge graph by detecting file languages, lazily loading Tree-sitter grammars, parsing source into ASTs, and walking the syntax tree to extract nodes and edges that represent code entities and their relationships.

The code-review-graph repository provides a framework for transforming raw source code into a queryable knowledge graph. By leveraging Tree-sitter's incremental parsing capabilities, it extracts semantic entities and their relationships from multiple programming languages without requiring language-specific compilers. This article examines how the parser transforms source files into nodes and edges using the Tree-sitter AST.

Language Detection via Extensions and Shebangs

Before parsing begins, CodeParser.detect_language identifies the programming language of each file.

Extension-Based Mapping

The parser consults the EXTENSION_TO_LANGUAGE lookup table defined in code_review_graph/parser.py (lines 13-78). This dictionary maps file extensions to their corresponding language identifiers, enabling instant classification of common file types like .py or .js.

Shebang Fallback for Extensionless Files

For files without extensions, the parser falls back to _detect_language_from_shebang (lines 21-88). This method inspects the shebang line (#!/usr/bin/env python3) to determine the appropriate language, ensuring scripts are correctly identified regardless of naming conventions.

Lazy Loading of Tree-sitter Parsers

Once the language is identified, CodeParser._get_parser manages the lifecycle of Tree-sitter grammar instances.

Per-Process Caching

The method first checks self._parsers, a per-process cache that avoids reloading grammars for files of the same language. This optimization reduces memory overhead and eliminates redundant initialization costs when processing large repositories.

Safe Grammar Loading with Subprocess Probing

If a grammar isn't cached, _load_tree_sitter_parser (lines 60-71) handles the loading. Before importing the native binary, _run_parser_load_probe (lines 75-88) executes a short-lived subprocess to verify the grammar can be loaded without crashing the main process. Upon successful verification, it imports tree_sitter_language_pack.get_parser (line 65) and stores the result in the cache.

AST Parsing and Entity Extraction

With a valid parser instance, the system converts source code into graph elements.

From Bytes to Syntax Tree

The parse_bytes method (lines 99-103 in parser.py) accepts raw file content and produces a Tree-sitter syntax tree:

parser = self._get_parser(language)          # Load the TS parser

tree = parser.parse(source)                  # Generate TS AST

Walking the Tree for Semantic Entities

The private method _walk_tree recursively visits every node in the Tree-sitter AST. It compares each node's type against language-specific tables including _CLASS_TYPES, _FUNCTION_TYPES, _IMPORT_TYPES, and _CALL_TYPES (defined in constants.py and referenced in parser.py lines 98-162).

For each matched node, the walker creates:

  • NodeInfo objects representing classes, functions, and types
  • EdgeInfo objects representing relationships like CALLS, IMPORTS_FROM, INHERITS, and CONTAINS

Dead Code Detection

During the walk, _is_in_static_dead_guard (lines 88-138) identifies code blocks that are statically unreachable (such as if False: constructs). This prevents the parser from adding spurious CALL edges to functions that are never actually invoked at runtime.

Graph Normalization and Assembly

Path Normalization

Both NodeInfo and EdgeInfo objects undergo path normalization via normalize_file_path (lines 50-66). This converts all file paths to POSIX-style separators, ensuring the knowledge graph remains portable across Windows and Unix-like systems.

Aggregating Results into the Graph

The parse_bytes method returns two Python lists: nodes and edges. These collections are passed to code_review_graph.graph.Graph, which merges per-file results into an in-memory adjacency structure. The Graph class handles deduplication and provides methods for persisting the structure to JSON or SQLite.

Repository-Wide Graph Construction

The top-level build_or_update_graph function in code_review_graph/tools/build.py (line 465) orchestrates the complete pipeline. It iterates over every source file in the repository, invokes CodeParser for each, and accumulates the results into a single Graph instance. Once complete, the graph can be queried via the CLI (code_review_graph/cli.py) or exported for visualization.

Code Examples

Parsing Individual Files

from pathlib import Path
from code_review_graph.parser import CodeParser

# Create a parser for a repo (optional repo_root for custom languages)

parser = CodeParser(repo_root=Path("/my/project"))

# Parse a single file – returns lists of NodeInfo and EdgeInfo objects

file_path = Path("src/example.py")
nodes, edges = parser.parse_file(file_path)

print("Nodes extracted:")
for n in nodes:
    print(f"{n.kind} {n.name} ({n.file_path}:{n.line_start})")

print("\nEdges extracted:")
for e in edges:
    print(f"{e.kind} {e.source} → {e.target} (line {e.line})")

Building the Complete Graph Programmatically

from code_review_graph.graph import Graph
from code_review_graph.parser import CodeParser
import os

parser = CodeParser()
graph = Graph()

for root, _, files in os.walk("."):
    for fn in files:
        p = Path(root, fn)
        if parser.detect_language(p) is None:
            continue
        nodes, edges = parser.parse_file(p)
        graph.add_nodes(nodes)
        graph.add_edges(edges)

graph.dump("graph.json")   # Persist for later queries / visualization

Command Line Usage


# Build the entire graph for the current repository

crg build            # Equivalent to code_review_graph/tools/build.py

Summary

  • Language Detection: CodeParser.detect_language uses extension mapping (EXTENSION_TO_LANGUAGE) and shebang inspection (_detect_language_from_shebang) to classify files.
  • Parser Management: _get_parser implements lazy loading with per-process caching and subprocess probing (_run_parser_load_probe) to safely initialize Tree-sitter grammars.
  • AST Extraction: parse_bytes generates Tree-sitter trees, while _walk_tree extracts NodeInfo and EdgeInfo objects using language-specific node type tables.
  • Graph Construction: The Graph class aggregates normalized nodes and edges, deduplicating entities across files and supporting export to JSON/SQLite.
  • Entry Point: build_or_update_graph in tools/build.py coordinates full-repository scanning and graph assembly.

Frequently Asked Questions

What is the role of Tree-sitter in code-review-graph?

Tree-sitter provides the incremental parsing engine that converts source code into concrete syntax trees. The code-review-graph parser uses these ASTs to identify classes, functions, imports, and calls without requiring language-specific compilers or static analysis tools, enabling support for multiple languages through a unified interface.

How does code-review-graph handle language detection for files without extensions?

For extensionless files, the parser executes _detect_language_from_shebang, which reads the shebang line (e.g., #!/usr/bin/env python3) to determine the interpreter. This allows the system to correctly identify and parse script files that lack standard file extensions.

Why does the parser use a subprocess probe before loading Tree-sitter grammars?

The _run_parser_load_probe method creates a short-lived subprocess to test-load the native Tree-sitter grammar binary. This isolation prevents corrupted or incompatible grammar files from crashing the main parsing process, ensuring robust operation when processing repositories with potentially damaged language packs.

How are file paths normalized in the knowledge graph?

The normalize_file_path function (lines 50-66 in parser.py) converts all file paths to POSIX-style forward slashes during NodeInfo and EdgeInfo creation. This normalization ensures the graph is portable between Windows and Unix-like operating systems, eliminating path separator inconsistencies in the output data.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →