# How code-review-graph Builds Its Knowledge Graph from Source Code Using Tree-sitter

> Discover how code-review-graph builds its knowledge graph from source code with Tree-sitter. Learn about language detection, AST parsing, and relationship extraction for comprehensive code analysis.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: how-to-guide
- Published: 2026-08-15

---

**The `code-review-graph` tool constructs a language-agnostic knowledge graph by detecting file languages, lazily loading Tree-sitter grammars, parsing source into ASTs, and walking the syntax tree to extract nodes and edges that represent code entities and their relationships.**

The `code-review-graph` repository provides a framework for transforming raw source code into a queryable knowledge graph. By leveraging Tree-sitter's incremental parsing capabilities, it extracts semantic entities and their relationships from multiple programming languages without requiring language-specific compilers. This article examines how the parser transforms source files into nodes and edges using the Tree-sitter AST.

## Language Detection via Extensions and Shebangs

Before parsing begins, `CodeParser.detect_language` identifies the programming language of each file.

### Extension-Based Mapping

The parser consults the `EXTENSION_TO_LANGUAGE` lookup table defined in [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py) (lines 13-78). This dictionary maps file extensions to their corresponding language identifiers, enabling instant classification of common file types like `.py` or `.js`.

### Shebang Fallback for Extensionless Files

For files without extensions, the parser falls back to `_detect_language_from_shebang` (lines 21-88). This method inspects the shebang line (`#!/usr/bin/env python3`) to determine the appropriate language, ensuring scripts are correctly identified regardless of naming conventions.

## Lazy Loading of Tree-sitter Parsers

Once the language is identified, `CodeParser._get_parser` manages the lifecycle of Tree-sitter grammar instances.

### Per-Process Caching

The method first checks `self._parsers`, a per-process cache that avoids reloading grammars for files of the same language. This optimization reduces memory overhead and eliminates redundant initialization costs when processing large repositories.

### Safe Grammar Loading with Subprocess Probing

If a grammar isn't cached, `_load_tree_sitter_parser` (lines 60-71) handles the loading. Before importing the native binary, `_run_parser_load_probe` (lines 75-88) executes a short-lived subprocess to verify the grammar can be loaded without crashing the main process. Upon successful verification, it imports `tree_sitter_language_pack.get_parser` (line 65) and stores the result in the cache.

## AST Parsing and Entity Extraction

With a valid parser instance, the system converts source code into graph elements.

### From Bytes to Syntax Tree

The `parse_bytes` method (lines 99-103 in [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py)) accepts raw file content and produces a Tree-sitter syntax tree:

```python
parser = self._get_parser(language)          # Load the TS parser

tree = parser.parse(source)                  # Generate TS AST

```

### Walking the Tree for Semantic Entities

The private method `_walk_tree` recursively visits every node in the Tree-sitter AST. It compares each node's type against language-specific tables including `_CLASS_TYPES`, `_FUNCTION_TYPES`, `_IMPORT_TYPES`, and `_CALL_TYPES` (defined in [`constants.py`](https://github.com/tirth8205/code-review-graph/blob/main/constants.py) and referenced in [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py) lines 98-162).

For each matched node, the walker creates:

- **NodeInfo** objects representing classes, functions, and types
- **EdgeInfo** objects representing relationships like CALLS, IMPORTS_FROM, INHERITS, and CONTAINS

### Dead Code Detection

During the walk, `_is_in_static_dead_guard` (lines 88-138) identifies code blocks that are statically unreachable (such as `if False:` constructs). This prevents the parser from adding spurious CALL edges to functions that are never actually invoked at runtime.

## Graph Normalization and Assembly

### Path Normalization

Both `NodeInfo` and `EdgeInfo` objects undergo path normalization via `normalize_file_path` (lines 50-66). This converts all file paths to POSIX-style separators, ensuring the knowledge graph remains portable across Windows and Unix-like systems.

### Aggregating Results into the Graph

The `parse_bytes` method returns two Python lists: `nodes` and `edges`. These collections are passed to `code_review_graph.graph.Graph`, which merges per-file results into an in-memory adjacency structure. The `Graph` class handles deduplication and provides methods for persisting the structure to JSON or SQLite.

## Repository-Wide Graph Construction

The top-level `build_or_update_graph` function in [`code_review_graph/tools/build.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/tools/build.py) (line 465) orchestrates the complete pipeline. It iterates over every source file in the repository, invokes `CodeParser` for each, and accumulates the results into a single `Graph` instance. Once complete, the graph can be queried via the CLI ([`code_review_graph/cli.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/cli.py)) or exported for visualization.

## Code Examples

### Parsing Individual Files

```python
from pathlib import Path
from code_review_graph.parser import CodeParser

# Create a parser for a repo (optional repo_root for custom languages)

parser = CodeParser(repo_root=Path("/my/project"))

# Parse a single file – returns lists of NodeInfo and EdgeInfo objects

file_path = Path("src/example.py")
nodes, edges = parser.parse_file(file_path)

print("Nodes extracted:")
for n in nodes:
    print(f"{n.kind} {n.name} ({n.file_path}:{n.line_start})")

print("\nEdges extracted:")
for e in edges:
    print(f"{e.kind} {e.source} → {e.target} (line {e.line})")

```

### Building the Complete Graph Programmatically

```python
from code_review_graph.graph import Graph
from code_review_graph.parser import CodeParser
import os

parser = CodeParser()
graph = Graph()

for root, _, files in os.walk("."):
    for fn in files:
        p = Path(root, fn)
        if parser.detect_language(p) is None:
            continue
        nodes, edges = parser.parse_file(p)
        graph.add_nodes(nodes)
        graph.add_edges(edges)

graph.dump("graph.json")   # Persist for later queries / visualization

```

### Command Line Usage

```bash

# Build the entire graph for the current repository

crg build            # Equivalent to code_review_graph/tools/build.py

```

## Summary

- **Language Detection**: `CodeParser.detect_language` uses extension mapping (`EXTENSION_TO_LANGUAGE`) and shebang inspection (`_detect_language_from_shebang`) to classify files.
- **Parser Management**: `_get_parser` implements lazy loading with per-process caching and subprocess probing (`_run_parser_load_probe`) to safely initialize Tree-sitter grammars.
- **AST Extraction**: `parse_bytes` generates Tree-sitter trees, while `_walk_tree` extracts `NodeInfo` and `EdgeInfo` objects using language-specific node type tables.
- **Graph Construction**: The `Graph` class aggregates normalized nodes and edges, deduplicating entities across files and supporting export to JSON/SQLite.
- **Entry Point**: `build_or_update_graph` in [`tools/build.py`](https://github.com/tirth8205/code-review-graph/blob/main/tools/build.py) coordinates full-repository scanning and graph assembly.

## Frequently Asked Questions

### What is the role of Tree-sitter in code-review-graph?

Tree-sitter provides the incremental parsing engine that converts source code into concrete syntax trees. The `code-review-graph` parser uses these ASTs to identify classes, functions, imports, and calls without requiring language-specific compilers or static analysis tools, enabling support for multiple languages through a unified interface.

### How does code-review-graph handle language detection for files without extensions?

For extensionless files, the parser executes `_detect_language_from_shebang`, which reads the shebang line (e.g., `#!/usr/bin/env python3`) to determine the interpreter. This allows the system to correctly identify and parse script files that lack standard file extensions.

### Why does the parser use a subprocess probe before loading Tree-sitter grammars?

The `_run_parser_load_probe` method creates a short-lived subprocess to test-load the native Tree-sitter grammar binary. This isolation prevents corrupted or incompatible grammar files from crashing the main parsing process, ensuring robust operation when processing repositories with potentially damaged language packs.

### How are file paths normalized in the knowledge graph?

The `normalize_file_path` function (lines 50-66 in [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py)) converts all file paths to POSIX-style forward slashes during `NodeInfo` and `EdgeInfo` creation. This normalization ensures the graph is portable between Windows and Unix-like operating systems, eliminating path separator inconsistencies in the output data.